Article
7 minute readllms.txt vs robots.txt: What Each File Controls
llms.txt vs robots.txt explained: one file sets crawler access, the other recommends pages to AI tools. Learn how they interact and which to fix first.
The llms.txt vs robots.txt question comes up because both are plain text files at the root of a site and both mention bots. Yet they do opposite jobs. robots.txt tells crawlers what they may fetch. llms.txt, by contrast, suggests what is worth reading. Mixing the two up can block pages you want cited or expose paths you meant to hide.
This guide explains what each file controls, how they interact and which one to fix first. It also shows a side-by-side comparison you can share with developers. So by the end, you will know which file to edit for each goal, and which goals neither file can reach.
The short answer
robots.txt is an access rule. llms.txt is a reading list. If you want to stop a crawler from fetching a section, edit robots.txt. If you want to point AI tools to your best pages, edit llms.txt.
Neither file is a security control, and neither one guarantees how a search engine or assistant ranks or cites you. Keep both facts in mind, because most mistakes come from expecting one file to do the other's job.
What robots.txt does
robots.txt is defined in RFC 9309, the Robots Exclusion Protocol, published in September 2022. The file must live at /robots.txt on the host it governs. It groups allow and disallow rules by user agent, such as Googlebot, Bingbot or GPTBot.
Its scope is crawling, not indexing. Google's robots.txt introduction says the file is not a way to keep a page out of Google; for that, use noindex or password protection. The RFC also says the rules are not a form of access authorization.
Also note that some fetches ignore it by design. OpenAI's crawler documentation says OAI-SearchBot and GPTBot respect robots.txt. However, ChatGPT-User acts on a user's request, so robots.txt rules may not apply to it.
What llms.txt does
llms.txt is a proposal published at llmstxt.org by Jeremy Howard in September 2024 and revised as version 2 in August 2026. It is a Markdown file with a title, a short summary and sections of links with notes. The goal is to give language models a curated map instead of a full crawl.
It allows and blocks nothing. A tool that reads the file can still fetch any page robots.txt permits, and a tool that ignores the file loses nothing. Moreover, Google's AI optimization guide says Google Search does not need AI text files to show your pages.
llms.txt vs robots.txt side by side
The table compares the two files on the points that matter most when you plan changes. Use it as a quick reference when a stakeholder asks which file to edit.
| Question | robots.txt | llms.txt |
|---|---|---|
| Main job | Allow or disallow crawling | Point to the most useful pages |
| Status | IETF standard, RFC 9309 | Community proposal, version 2 |
| Format | User-agent groups and path rules | Markdown title, summary and link lists |
| Location | /robots.txt on each host | /llms.txt at the root or a subfolder |
| Effect if missing | Everything may be crawled | No effect on crawling |
| Used by Google Search | Yes, for crawling | No |
How the two files interact
The files live side by side, but robots.txt always wins on access. Three practical rules follow from that.
Blocked pages stay blocked
If robots.txt disallows /docs/ for a crawler, a link to /docs/api in llms.txt does not unlock it. A compliant crawler will read your reading list and then decline to fetch the page.
So before you add a link to llms.txt, test the URL against your robots.txt rules for the agents you care about. A link that no relevant bot may fetch only adds noise to the file.
Don't use llms.txt to hide content
Leaving a page out of llms.txt does not keep it away from AI tools. They can still find it through your sitemap, internal links or a search engine. If a page must stay private, protect it with authentication rather than with either text file.
Likewise, never list private paths in robots.txt to hide them. Both files are public, and anyone can read them. A disallow line for /internal-reports/ tells every visitor where to look.
Keep the two files consistent
Review both files together whenever you change site structure. For example, a migration from /help/ to /support/ needs new robots rules, new llms.txt links and redirects for the old paths.
A simple test catches most drift. Take every URL in llms.txt and confirm that it returns 200 and is allowed for the crawlers you target. Then fix or remove any link that fails.
Which file to fix first
If you have limited time, fix robots.txt first, because a crawl mistake affects every search engine and assistant at once. Then work through this order:
- Accidental blocks: make sure no important section is disallowed for search crawlers.
- Intentional AI rules: decide per agent whether training crawlers and search crawlers may fetch your content.
- Sitemap line: add a
Sitemap:line so crawlers find your canonical URLs. - llms.txt: publish a curated file once access rules are settled.
This order matters because llms.txt depends on access. A perfect reading list helps nobody if the pages behind it are blocked.
Example: llms.txt vs robots.txt on one site
Consider a fictional software company that wants search engines and AI search tools to read its docs, but does not want its content used for model training. It also has a staging area that should never appear anywhere. Here is a robots.txt that expresses those access decisions:
User-agent: GPTBot
Disallow: /
User-agent: *
Disallow: /cart/
Sitemap: https://www.example.com/sitemap.xmlNotice what is missing: the staging area. It sits behind a login instead, because a disallow line would advertise it. Meanwhile, OAI-SearchBot falls under the general group, so ChatGPT search can still fetch the docs.
The llms.txt file then lists only pages those rules allow, with notes that say when each one helps:
# Example Software
> Example Software makes scheduling tools for clinics.
## Docs
- [Getting started](https://www.example.com/docs/start): Set up a first calendar
- [API reference](https://www.example.com/docs/api): Endpoints and limitsThat split is the whole point of llms.txt vs robots.txt. Each file does one job: robots.txt draws the access boundary, while llms.txt highlights the best pages inside it.
Conclusion
In the llms.txt vs robots.txt comparison, robots.txt controls access and llms.txt recommends content. robots.txt is a standard that Google and other major crawlers follow. llms.txt is an optional proposal that some AI tools read, while Google Search does not use it.
Therefore, start with robots.txt, confirm that your important pages are crawlable, then add an llms.txt file that points to them. Review both whenever your URLs change, and protect anything private with authentication instead.
Sources
These sources were reviewed on October 7, 2026. The comparison and recommendations are independently written by SEOVision.
- RFC 9309: Robots Exclusion Protocol: location, rule syntax and limits of robots.txt
- Google Search Central: Introduction to robots.txt: what robots.txt does and does not do
- OpenAI: Overview of OpenAI crawlers: which agents respect robots.txt
- The /llms.txt file proposal, version 2: format and purpose of llms.txt
- Google Search Central: AI optimization guide: Google's position on AI text files
Frequently asked questions
Quick answers to the questions readers ask most about this topic.
What is the difference between llms.txt and robots.txt?
robots.txt tells crawlers which URLs they may fetch. llms.txt is a Markdown reading list that points AI tools to your most useful pages; it allows or blocks nothing.
Can llms.txt override robots.txt?
No. If robots.txt disallows a URL for a crawler, listing that URL in llms.txt does not make it fetchable for a compliant crawler.
Do I need both llms.txt and robots.txt?
robots.txt is worth having on almost every site, if only for the Sitemap line. llms.txt is optional; it costs little, but Google Search does not use it.
Can I block AI crawlers with llms.txt?
No. Use robots.txt rules for each AI user agent, such as GPTBot. Note that some user-initiated fetches, like ChatGPT-User, may not follow robots.txt.
Sources
These references support the platform guidance discussed above. Worked examples are illustrative unless identified as measured results.
Keep reading