AI crawlers: block the training bots, feed the ones that cite you

Block AI crawlers the right way: stop the training-only bots like GPTBot, but let the answer bots cite you. Why llms.txt won't help and robots.txt isn't a wall.
Block AI crawlers the right way: stop the training-only bots like GPTBot, but let the answer bots cite you. Why llms.txt won't help and robots.txt isn't a wall.
A 12-point website conversion audit founders can run in 30 minutes. Score yes/no, fix the no's first, and stop redesigning before diagnosing.

Gabriel Espinheira

Blocking every AI crawler feels like the safe move. For an owner-operated business, it is usually the expensive one. The scare-pieces tell you the bots are stealing your content, so you paste a block-everything snippet into robots.txt and move on. Without meaning to, you switch off the one channel that is growing while Google search quietly shrinks.

The decision that actually matters is not whether to block AI crawlers, but which ones. Some bots exist to train models on your work. Others exist to fetch a page and cite you inside an answer a buyer is reading right now. Treat them the same and you either give away your content for nothing or vanish from the tool your customers increasingly ask first. Separate the two and it becomes a two-line robots.txt decision you can make in ten minutes.

TL;DR: Don't block AI crawlers wholesale. Block the training-only bots (GPTBot, ClaudeBot, Google-Extended, CCBot, Meta-ExternalAgent) if protecting your content matters to you; it does not touch your Google rankings. Always allow the search and answer bots (OAI-SearchBot, PerplexityBot, ChatGPT-User, Claude-User) so AI tools can cite you. And skip llms.txt: no major engine reads it yet.

Blocking a crawler and blocking a citation are two different decisions

An AI crawler is not one thing, and that is the whole game. OpenAI alone runs several. GPTBot pulls pages to train models, OAI-SearchBot fetches pages so ChatGPT can cite them in search answers, and ChatGPT-User acts when a person clicks a link inside a chat. Anthropic, Google, and Perplexity each split their bots the same way: one lane for training, one lane for answering.

That split is the decision. Blocking a training bot opts your content out of the next model, which is a genuine intellectual-property choice. Blocking a search or answer bot removes you from the AI result a buyer is looking at, which is a visibility choice and almost always the wrong one. The listicles that hand you a block-everything snippet collapse those two decisions into one checkbox. Google's own crawler documentation and OpenAI's both make the separation explicit: block GPTBot and your Google rankings do not move an inch, because Googlebot is a different bot entirely.

The traffic you are "protecting" is not the traffic you get

Look at what the training bots actually cost you, and the protection instinct starts to make sense. In the week of 19 to 26 June 2025, Cloudflare measured Anthropic's crawler making roughly 71,000 page requests for every single visitor it sent back. Google's crawler, by comparison, has always returned a visitor for every handful of pages it takes. By that measure, blocking the training bots looks obvious: they read everything and give nothing.

That is only half the ledger, and it is the cheap half. The visitor an AI tool does send is worth more than the ones you are counting. Pew Research Center's 2025 study of real US browsing found people clicked a traditional link on just 8% of searches that showed an AI summary, against 15% when none appeared, so the click you used to get is already leaking away whether you block anything or not. Meanwhile several 2026 analyses put the conversion rate of AI-referred visitors well above Google organic: someone who arrives because ChatGPT recommended you has already been pre-sold by the answer. You are not protecting a vault. You are boarding up the window your next customer looks through.

Disappearing from the answer box costs more than a scrape

The reason this matters now, and not in two years, is that the answer box is quietly becoming the listing. When Google shows an AI Overview, the blue links below it stop earning clicks. Seer Interactive's November 2025 analysis of 42 organisations found organic click-through on AI Overview queries fell 61% between June 2024 and September 2025. Ahrefs, studying 300,000 keywords, put the drop on the top result at about 34.5% when an Overview appears.

For an owner-operated European business, that is the real exposure. Not that a model trained on your blog, but that a founder searching your category gets a synthesised answer naming three businesses, and you are not one of them. Block the search and answer bots and you are not eligible to be one of them. The scrape you were afraid of costs you nothing you can measure. The invisibility costs you the enquiry.

llms.txt is not the shortcut everyone is selling

You will be told the answer is an llms.txt file: a tidy list of your best pages, dropped at the root of your site, that supposedly tells AI models what to read. Don't bother yet. Google has said plainly it does not use it. John Mueller compared llms.txt to the old keywords meta tag that search engines ignored for a decade, and in July 2025 Gary Illyes confirmed Google has no plans to support it. As of early 2026, OpenAI, Anthropic, and Meta have not committed to reading it in production either, and OpenAI's crawler docs still tell you to control its bots with robots.txt.

So llms.txt is a file the reputable engines do not read, sold as the fix for a problem robots.txt already handles. Write one if it costs you five minutes and you like tidy things. Do not treat it as visibility work, and do not let anyone bill you for it as a growth service.

robots.txt is a request, not a wall

Here is the part the how-to guides leave out: even the lever that does work is not enforcement. robots.txt is a note on the door asking bots to behave. It is not a lock. In August 2025 Cloudflare delisted Perplexity as a verified crawler, accusing it of using undeclared bots and rotating user agents and networks to fetch pages from sites that had explicitly blocked it. Perplexity denied it. Whatever the verdict, the lesson stands: a determined crawler can ignore your file.

That does not make robots.txt pointless. The reputable bots, the ones you actually want a relationship with, honour it, which is exactly why the training-versus-answer split is the whole decision. It does mean you should stop treating "I blocked them" as "they are gone." If a page genuinely cannot be public, such as pricing you negotiate privately or client material under an agreement, the answer is authentication or keeping it off the public site, not a line in a text file you are trusting strangers to read.

The two-line decision for a European founder

Strip away the noise and this is a short, reversible robots.txt edit, not an existential one. Make it once and move on.

  • If protecting your content matters to you, block the training-only bots. Add Disallow rules for GPTBot, ClaudeBot, Google-Extended, CCBot, and Meta-ExternalAgent. This opts you out of model training and does not affect your Google Search rankings or your AI Overview eligibility.

  • Always allow the search and answer bots. Leave OAI-SearchBot, PerplexityBot, ChatGPT-User, and Claude-User free to fetch, so AI tools can quote and cite you when a buyer asks about your category.

  • Then do the work that actually earns the citation. Being crawlable is permission, not performance. A clear direct answer near the top of each page, specific numbers, and named sources are what get you quoted.

If you only ever change one thing, make it the middle line. The default block-everything snippet gets that one backwards, and it is the one that costs you enquiries.

Frequently asked questions

Should I block AI crawlers on my website?

Block the training-only bots (GPTBot, ClaudeBot, Google-Extended, CCBot, Meta-ExternalAgent) if opting out of AI model training matters to you. Do not block the search and answer bots. Blocking those removes you from the AI results your buyers now read before they ever reach a blue link.

Does blocking GPTBot or ClaudeBot hurt my Google rankings?

No. GPTBot and ClaudeBot are separate from Googlebot, and both Google's and OpenAI's own documentation confirm it. Blocking them opts you out of AI training with zero effect on your Google Search position or your eligibility to appear in an AI Overview.

Does an llms.txt file help AI find my content?

Not yet. Google has publicly said it does not read llms.txt, and as of early 2026 OpenAI, Anthropic, and Meta have not committed to it in production. Control access with robots.txt instead, which the reputable crawlers actually honour.

Can robots.txt actually stop AI crawlers?

It stops the ones that choose to obey it, which includes the major reputable crawlers, but it is a request rather than a hard block. Cloudflare accused Perplexity of evading blocks in 2025. For content that truly cannot be public, use authentication, not a text file.

Blocking AI crawlers is not the shield it looks like, and llms.txt is not the fix it is sold as. The move that protects you without erasing you is narrow: block the bots that only take, feed the bots that cite, and spend the saved effort making each page worth quoting. robots.txt is a request, not a wall, and llms.txt is not even a request anyone is reading.

Plan. Build. Iterate. Want an honest read on whether AI is sending you traffic, or sending it to your competitors? Book a 30-min call and we will look at your setup together. Or see how SharpOS, the workspace inside every SharpHaw subscription, shows the work in motion every week.

Ready to start?

Book a 30-minute call. We'll dig into what's working, what isn't, and what the first move should be. No fluff, no pressure. If it makes sense to work together, we'll make it happen.

Ready to start?

Book a 30-minute call. We'll dig into what's working, what isn't, and what the first move should be. No fluff, no pressure. If it makes sense to work together, we'll make it happen.

Read more