Insights — Digital Infrastructure — 5 min read
Should You Allow or Block AI Crawlers? GPTBot, robots.txt and llms.txt Explained
Training crawlers and search crawlers are not the same thing. Blocking the wrong one can quietly remove a business from ChatGPT or Copilot without anyone noticing.

In short
There is no single right answer for every business, because 'AI crawler' covers training bots (like GPTBot and Google-Extended) and search bots (like OAI-SearchBot, PerplexityBot and Bingbot) with different purposes. Blocking training crawlers has no effect on whether ChatGPT, Gemini or Copilot can surface your pages in an answer; blocking the relevant search crawler generally removes you from that platform's answers entirely. Decide deliberately, crawler by crawler, rather than applying one blanket rule.
Many businesses have never looked at their robots.txt file, and some that have added blanket 'block all AI' rules — often copied from a template — without realising that 'AI crawler' covers several different bots doing different jobs. One trains a model. Another finds pages for a live search answer. Blocking all of them with one rule can switch off a channel a company intended to keep open.
This article sets out which crawlers do what, how to check what your own site currently allows, and the genuine, limited evidence on llms.txt.
Training crawlers and search crawlers are not the same thing
The most common misunderstanding is treating 'AI crawler' as one category. In practice, the major providers separate crawlers used to gather data for training a model from crawlers used to find and surface pages in a live search or answer feature. Blocking a training crawler is a decision about whether your content can be used to improve a model. Blocking a search crawler is a decision about whether your pages can appear in that platform's answers at all. Confusing the two is the single most common and consequential mistake.
The crawlers that matter, and what each one does
| Crawler | Operator | Purpose | Blocking it means |
|---|---|---|---|
| GPTBot | OpenAI | Model training | Content excluded from future training data; does not affect ChatGPT search answers |
| OAI-SearchBot | OpenAI | Finding pages for ChatGPT's search feature | Site will not appear in ChatGPT search answers (may still appear as a plain link) |
| ChatGPT-User | OpenAI | Fetching a page a user references in a live chat | That on-demand fetch fails; unrelated to broader search inclusion |
| Google-Extended | Gemini model training/grounding | Content excluded from Gemini training/grounding use; does not affect Google Search indexing or ranking | |
| Googlebot | Standard Search indexing | Site disappears from Google Search entirely, including AI Overviews/AI Mode, which depend on normal Search eligibility | |
| PerplexityBot | Perplexity | Indexing for Perplexity's search answers | Site excluded from Perplexity's indexed answers |
| Perplexity-User | Perplexity | On-demand fetch triggered by a user | That specific fetch fails |
| Bingbot | Microsoft | Standard Bing indexing | Site disappears from Bing, and from Copilot, which Microsoft documents as grounding its web answers on Bing |
Should you block anything at all?
For most B2B companies seeking visibility and enquiries, the default of allowing all of these crawlers is the sensible starting point, because blocking them removes the possibility of appearing in the corresponding answers without any guarantee of a compensating benefit. There are legitimate reasons a business might still choose to block a training crawler specifically: concerns about proprietary pricing, technical documentation or client-identifiable material being used to train a third-party model. That is a defensible commercial decision — distinct from blocking a search crawler, which works directly against a visibility goal.
How to check what your website currently allows
- 01Open yourdomain.com/robots.txt directly in a browser and read every 'Disallow' line against the crawler list above
- 02Check whether a CDN, firewall or bot-protection service (e.g. Cloudflare, a WAF, or a hosting platform's bot rules) is blocking these user agents at the network level — robots.txt can look permissive while a separate service silently blocks the request
- 03Check server or CDN access logs for the crawler names above to see whether they have actually been visiting, not just whether they are theoretically allowed
- 04In Google Search Console and Bing Webmaster Tools, confirm key pages show as indexed, which rules out the most common underlying cause
- 05If using a website platform or agency-managed site, ask explicitly whether any 'security' or 'bot blocking' feature has been enabled that affects these crawlers, since this is often switched on by default without discussion
Why Google can find a site but ChatGPT cannot
This is one of the most common support questions, and it usually comes down to one of three causes: OAI-SearchBot specifically blocked in robots.txt while Googlebot is allowed; a bot-protection or firewall service blocking OpenAI's user agents by default while allowing Google's (which many services treat as 'safe' by reputation); or content that only renders after client-side JavaScript runs, which Google processes (in a secondary rendering pass) more reliably than some AI crawlers are reported to. None of these is a penalty — they are configuration differences that are straightforward to check and fix.
Can AI crawlers read JavaScript-heavy websites?
Google documents that it renders JavaScript, but processes it in stages, and recommends not relying solely on client-side rendering for important content. It is widely reported — though not something every provider documents in full technical detail — that several AI-related crawlers do not reliably execute JavaScript, meaning content that only appears after a script runs may be invisible to them even though a human visitor sees it fine. The safest practice is ensuring core facts (what the company does, for whom, key services, contact details) exist in the server-rendered HTML, not only in a client-side-rendered component.
What is llms.txt, and do you need one?
llms.txt is a community-proposed file, similar in spirit to robots.txt or a sitemap, intended to give AI systems a concise summary of a site's key pages. It is not an official standard from any AI provider. As of current public information, no major AI search platform has publicly committed to reading or using it, and Google has stated it does not use it for Search. Creating one is low-cost and causes no harm, but there is no publicly documented evidence it improves AI visibility. Treat it as a minor, optional addition — not a priority ahead of crawlability, clear content and indexing fundamentals.
A simple decision framework
- Visibility is the goal: allow GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Google-Extended, Googlebot and Bingbot unless there is a specific reason not to
- Concerned specifically about model training on proprietary content: consider blocking GPTBot and Google-Extended only, leaving the search-facing crawlers (OAI-SearchBot, PerplexityBot, Googlebot, Bingbot) allowed
- Not sure what is currently configured: check robots.txt, bot-protection settings and server logs before assuming either way
- Site is JavaScript-heavy: confirm core facts exist in server-rendered HTML regardless of the robots.txt decision
Mistakes to avoid
- Copying a blanket 'block all AI bots' robots.txt template without checking which named crawlers it actually lists
- Assuming a CDN or security plugin's default settings allow these crawlers without checking
- Treating llms.txt as a substitute for being properly indexed by Google and Bing
- Reviewing robots.txt once and never again — platforms add and rename crawlers, and the right configuration is not fixed forever
Not sure why the website isn't being found or isn't producing enquiries?
A Commercial Website Audit reviews search visibility, structure, AI readiness and conversion together, and gives you a prioritised 90-day plan. £595 + VAT, fixed fee.
Related services
