Short answer: different crawlers visit sites for different purposes: indexing for search, gathering training data for AI models, and fetching pages live to answer a user’s question. robots.txt lets you allow or disallow crawlers by name, and well-behaved crawlers respect it, though it is a request rather than a technical barrier. Blocking the crawler that powers a search engine’s results also removes you from them. llms.txt is a proposed file offering AI systems a curated summary of a site; it is not an established standard and support varies. Decide based on what you want, and check what each rule actually blocks.
Website logs now show a wider range of automated visitors than a few years ago. Some index pages for search, some collect text for training AI models, and some fetch a page in real time because a user asked an assistant a question.
Site owners have some control over these visitors, and the choices involve genuine trade-offs. It helps to understand what each control does before changing anything.
Three kinds of automated visitor
The crawlers that matter for most sites fall into three groups.
| Tip | Purpose | Effect of blocking |
|---|---|---|
| Search crawlers | Index pages for search results | Pages disappear from that search engine |
| Training crawlers | Collect text to train AI models | Content not used for future training |
| User-triggered fetchers | Retrieve a page to answer a user’s question | Page not read for that answer |
Companies usually document their crawlers’ names and purposes. Check the current documentation before writing rules, as names and behaviours change.
What robots.txt can do
robots.txt is a plain text file at the root of your site listing rules for crawlers by name: which paths each may or may not access.
Well-behaved crawlers read and follow it. It is not access control: a crawler that ignores it can still fetch pages, and the file itself is public. For content that must not be accessed, use authentication, not robots.txt.
It also controls crawling, not indexing. A page blocked in robots.txt can still appear in results if other sites link to it; a noindex tag is the tool for keeping a page out of results.
The search trade-off
Blocking a search engine’s main crawler removes your pages from its results, and from features built on those results. That is almost never what a business blog wants.
Some companies offer separate controls for AI uses that are distinct from search indexing. Where they exist, they let you opt out of certain uses while remaining in search. Read the current documentation for each, because the details matter and change.
Deciding what to allow
The decision depends on what the site is for.
- A business blog aiming for visibility usually wants to be found and cited, and allows search crawlers and often user-triggered fetchers.
- A publisher whose content is the product may reasonably block training crawlers while allowing search.
- A site with private or licensed material needs authentication, not robots.txt.
There is no universally correct setting. Write down what you want and choose rules that match it.
What llms.txt is
llms.txt is a proposed convention: a text file at the root of a site giving AI systems a concise, curated overview of the site’s most important content, typically in simple markdown with links.
It is a proposal rather than an established standard, and support among AI systems varies and is not well documented. Creating one does little harm and may help some tools, but it should not be expected to change visibility on its own, and it does not replace a clear, well-structured site.
If you create an llms.txt
Keep it short and accurate: a one-paragraph description of the site, then links to the pages that best represent it, grouped sensibly — main guides, product information, pricing, contact.
Keep it in sync with the site. A file listing outdated pages or old prices does more harm than none. Treat it like a sitemap for people and tools, maintained when the site changes.
Checking what you have
Open yoursite.com/robots.txt and read it. Many sites have rules left over from development, or added by plugins, that block more than intended. why posts do not rank lists an over-broad robots.txt among the first things to rule out when pages do not appear.
Search Console’s URL inspection shows whether a page is blocked for its crawler. For other crawlers, server logs show which ones visit and what they request.
A simple starting configuration
For a typical business blog that wants to be visible, a reasonable starting point is a robots.txt that allows search crawlers everywhere except administrative paths, points to the sitemap, and makes a deliberate, documented choice about training crawlers. Review it when your priorities change or when crawler documentation changes.
Keep a note of why each rule exists. Rules added without explanation tend to outlive their reason and quietly block things later.
Server load
Some site owners notice heavy crawling from automated visitors. If crawling affects performance, rate limits at the server or hosting level are usually more effective than robots.txt, and your host can advise on them.
Keeping perspective
For most business blogs, the practical priority is simple: make sure search crawlers can reach the pages you want found, and make those pages clear and useful. The finer choices about training crawlers and llms.txt matter more for publishers whose content is itself the product.
None of these controls guarantees inclusion or citation in AI answers. They let you decide what you are willing to offer; the content decides whether it is used.
Related reading
If this was useful, these cover the questions that usually come next.
- AI answers in search and traffic — why these choices matter
- Setting up a new blog — the robots.txt basics
- Why posts do not rank — when robots.txt blocks too much
The bottom line
Know the three kinds of crawler: search, training and user-triggered. robots.txt asks crawlers to stay away from paths, is followed by well-behaved ones, and blocking a search crawler removes you from that search. llms.txt is an optional proposal, not a standard. Decide what you want, check current documentation, and make sure you are not blocking more than you meant to.
FAQ
What are AI crawlers?
Automated visitors that fetch web pages for AI purposes, such as collecting training data or retrieving pages to answer a user’s question, alongside traditional search crawlers.
Can robots.txt block AI crawlers?
It can ask named crawlers not to access paths, and well-behaved crawlers comply. It is a request, not access control.
Will blocking AI crawlers hurt my search visibility?
Blocking a search engine’s main crawler removes you from its results. Some companies provide separate controls for AI uses; check their current documentation.
What is llms.txt?
A proposed file at a site’s root that gives AI systems a curated overview of important content. It is not an established standard and support varies.
Should I create an llms.txt file?
It does little harm if kept short and accurate, but do not expect it to change visibility by itself. A clear, useful site matters more.
How do I check what my robots.txt blocks?
Read the file at yoursite.com/robots.txt, and use URL inspection in Search Console to check specific pages.


