Article

7 minute read

llms.txt vs robots.txt: What Each File Controls

llms.txt vs robots.txt explained: one file sets crawler access, the other recommends pages to AI tools. Learn how they interact and which to fix first.

Abstract illustration for llms.txt vs robots.txt: What Each File Controls

The llms.txt vs robots.txt question comes up because both are plain text files at the root of a site and both mention bots. Yet they do opposite jobs. robots.txt tells crawlers what they may fetch. llms.txt, by contrast, suggests what is worth reading. Mixing the two up can block pages you want cited or expose paths you meant to hide.

This guide explains what each file controls, how they interact and which one to fix first. It also shows a side-by-side comparison you can share with developers. So by the end, you will know which file to edit for each goal, and which goals neither file can reach.

The short answer

robots.txt is an access rule. llms.txt is a reading list. If you want to stop a crawler from fetching a section, edit robots.txt. If you want to point AI tools to your best pages, edit llms.txt.

Neither file is a security control, and neither one guarantees how a search engine or assistant ranks or cites you. Keep both facts in mind, because most mistakes come from expecting one file to do the other's job.

What robots.txt does

robots.txt is defined in RFC 9309, the Robots Exclusion Protocol, published in September 2022. The file must live at /robots.txt on the host it governs. It groups allow and disallow rules by user agent, such as Googlebot, Bingbot or GPTBot.

Its scope is crawling, not indexing. Google's robots.txt introduction says the file is not a way to keep a page out of Google; for that, use noindex or password protection. The RFC also says the rules are not a form of access authorization.

Also note that some fetches ignore it by design. OpenAI's crawler documentation says OAI-SearchBot and GPTBot respect robots.txt. However, ChatGPT-User acts on a user's request, so robots.txt rules may not apply to it.

What llms.txt does

llms.txt is a proposal published at llmstxt.org by Jeremy Howard in September 2024 and revised as version 2 in August 2026. It is a Markdown file with a title, a short summary and sections of links with notes. The goal is to give language models a curated map instead of a full crawl.

It allows and blocks nothing. A tool that reads the file can still fetch any page robots.txt permits, and a tool that ignores the file loses nothing. Moreover, Google's AI optimization guide says Google Search does not need AI text files to show your pages.

llms.txt vs robots.txt side by side

The table compares the two files on the points that matter most when you plan changes. Use it as a quick reference when a stakeholder asks which file to edit.

Questionrobots.txtllms.txt
Main jobAllow or disallow crawlingPoint to the most useful pages
StatusIETF standard, RFC 9309Community proposal, version 2
FormatUser-agent groups and path rulesMarkdown title, summary and link lists
Location/robots.txt on each host/llms.txt at the root or a subfolder
Effect if missingEverything may be crawledNo effect on crawling
Used by Google SearchYes, for crawlingNo

How the two files interact

The files live side by side, but robots.txt always wins on access. Three practical rules follow from that.

Blocked pages stay blocked

If robots.txt disallows /docs/ for a crawler, a link to /docs/api in llms.txt does not unlock it. A compliant crawler will read your reading list and then decline to fetch the page.

So before you add a link to llms.txt, test the URL against your robots.txt rules for the agents you care about. A link that no relevant bot may fetch only adds noise to the file.

Don't use llms.txt to hide content

Leaving a page out of llms.txt does not keep it away from AI tools. They can still find it through your sitemap, internal links or a search engine. If a page must stay private, protect it with authentication rather than with either text file.

Likewise, never list private paths in robots.txt to hide them. Both files are public, and anyone can read them. A disallow line for /internal-reports/ tells every visitor where to look.

Keep the two files consistent

Review both files together whenever you change site structure. For example, a migration from /help/ to /support/ needs new robots rules, new llms.txt links and redirects for the old paths.

A simple test catches most drift. Take every URL in llms.txt and confirm that it returns 200 and is allowed for the crawlers you target. Then fix or remove any link that fails.

Which file to fix first

If you have limited time, fix robots.txt first, because a crawl mistake affects every search engine and assistant at once. Then work through this order:

  • Accidental blocks: make sure no important section is disallowed for search crawlers.
  • Intentional AI rules: decide per agent whether training crawlers and search crawlers may fetch your content.
  • Sitemap line: add a Sitemap: line so crawlers find your canonical URLs.
  • llms.txt: publish a curated file once access rules are settled.

This order matters because llms.txt depends on access. A perfect reading list helps nobody if the pages behind it are blocked.

Example: llms.txt vs robots.txt on one site

Consider a fictional software company that wants search engines and AI search tools to read its docs, but does not want its content used for model training. It also has a staging area that should never appear anywhere. Here is a robots.txt that expresses those access decisions:

User-agent: GPTBot
Disallow: /

User-agent: *
Disallow: /cart/

Sitemap: https://www.example.com/sitemap.xml

Notice what is missing: the staging area. It sits behind a login instead, because a disallow line would advertise it. Meanwhile, OAI-SearchBot falls under the general group, so ChatGPT search can still fetch the docs.

The llms.txt file then lists only pages those rules allow, with notes that say when each one helps:

# Example Software

> Example Software makes scheduling tools for clinics.

## Docs

- [Getting started](https://www.example.com/docs/start): Set up a first calendar
- [API reference](https://www.example.com/docs/api): Endpoints and limits

That split is the whole point of llms.txt vs robots.txt. Each file does one job: robots.txt draws the access boundary, while llms.txt highlights the best pages inside it.

Conclusion

In the llms.txt vs robots.txt comparison, robots.txt controls access and llms.txt recommends content. robots.txt is a standard that Google and other major crawlers follow. llms.txt is an optional proposal that some AI tools read, while Google Search does not use it.

Therefore, start with robots.txt, confirm that your important pages are crawlable, then add an llms.txt file that points to them. Review both whenever your URLs change, and protect anything private with authentication instead.

Sources

These sources were reviewed on October 7, 2026. The comparison and recommendations are independently written by SEOVision.

Frequently asked questions

Quick answers to the questions readers ask most about this topic.

What is the difference between llms.txt and robots.txt?

robots.txt tells crawlers which URLs they may fetch. llms.txt is a Markdown reading list that points AI tools to your most useful pages; it allows or blocks nothing.

Can llms.txt override robots.txt?

No. If robots.txt disallows a URL for a crawler, listing that URL in llms.txt does not make it fetchable for a compliant crawler.

Do I need both llms.txt and robots.txt?

robots.txt is worth having on almost every site, if only for the Sitemap line. llms.txt is optional; it costs little, but Google Search does not use it.

Can I block AI crawlers with llms.txt?

No. Use robots.txt rules for each AI user agent, such as GPTBot. Note that some user-initiated fetches, like ChatGPT-User, may not follow robots.txt.

Sources

These references support the platform guidance discussed above. Worked examples are illustrative unless identified as measured results.

  1. RFC 9309 rfc-editor.org
  2. robots.txt introduction developers.google.com
  3. crawler documentation developers.openai.com
  4. llmstxt.org llmstxt.org
  5. AI optimization guide developers.google.com