Skip to content
Evans Sales Consultancy - international sales growth, market entry and expansionEvansSales Consultancy
Call 0330 043 8477Email

Insights — Digital Infrastructure — 5 min read

Should You Allow or Block AI Crawlers? GPTBot, robots.txt and llms.txt Explained

Training crawlers and search crawlers are not the same thing. Blocking the wrong one can quietly remove a business from ChatGPT or Copilot without anyone noticing.

A website owner checking robots.txt settings that control which AI crawlers can access their site

In short

There is no single right answer for every business, because 'AI crawler' covers training bots (like GPTBot and Google-Extended) and search bots (like OAI-SearchBot, PerplexityBot and Bingbot) with different purposes. Blocking training crawlers has no effect on whether ChatGPT, Gemini or Copilot can surface your pages in an answer; blocking the relevant search crawler generally removes you from that platform's answers entirely. Decide deliberately, crawler by crawler, rather than applying one blanket rule.

Many businesses have never looked at their robots.txt file, and some that have added blanket 'block all AI' rules — often copied from a template — without realising that 'AI crawler' covers several different bots doing different jobs. One trains a model. Another finds pages for a live search answer. Blocking all of them with one rule can switch off a channel a company intended to keep open.

This article sets out which crawlers do what, how to check what your own site currently allows, and the genuine, limited evidence on llms.txt.

Training crawlers and search crawlers are not the same thing

The most common misunderstanding is treating 'AI crawler' as one category. In practice, the major providers separate crawlers used to gather data for training a model from crawlers used to find and surface pages in a live search or answer feature. Blocking a training crawler is a decision about whether your content can be used to improve a model. Blocking a search crawler is a decision about whether your pages can appear in that platform's answers at all. Confusing the two is the single most common and consequential mistake.

The crawlers that matter, and what each one does

CrawlerOperatorPurposeBlocking it means
GPTBotOpenAIModel trainingContent excluded from future training data; does not affect ChatGPT search answers
OAI-SearchBotOpenAIFinding pages for ChatGPT's search featureSite will not appear in ChatGPT search answers (may still appear as a plain link)
ChatGPT-UserOpenAIFetching a page a user references in a live chatThat on-demand fetch fails; unrelated to broader search inclusion
Google-ExtendedGoogleGemini model training/groundingContent excluded from Gemini training/grounding use; does not affect Google Search indexing or ranking
GooglebotGoogleStandard Search indexingSite disappears from Google Search entirely, including AI Overviews/AI Mode, which depend on normal Search eligibility
PerplexityBotPerplexityIndexing for Perplexity's search answersSite excluded from Perplexity's indexed answers
Perplexity-UserPerplexityOn-demand fetch triggered by a userThat specific fetch fails
BingbotMicrosoftStandard Bing indexingSite disappears from Bing, and from Copilot, which Microsoft documents as grounding its web answers on Bing
Documented AI and search crawlers relevant to B2B visibility

Should you block anything at all?

For most B2B companies seeking visibility and enquiries, the default of allowing all of these crawlers is the sensible starting point, because blocking them removes the possibility of appearing in the corresponding answers without any guarantee of a compensating benefit. There are legitimate reasons a business might still choose to block a training crawler specifically: concerns about proprietary pricing, technical documentation or client-identifiable material being used to train a third-party model. That is a defensible commercial decision — distinct from blocking a search crawler, which works directly against a visibility goal.

How to check what your website currently allows

  1. 01Open yourdomain.com/robots.txt directly in a browser and read every 'Disallow' line against the crawler list above
  2. 02Check whether a CDN, firewall or bot-protection service (e.g. Cloudflare, a WAF, or a hosting platform's bot rules) is blocking these user agents at the network level — robots.txt can look permissive while a separate service silently blocks the request
  3. 03Check server or CDN access logs for the crawler names above to see whether they have actually been visiting, not just whether they are theoretically allowed
  4. 04In Google Search Console and Bing Webmaster Tools, confirm key pages show as indexed, which rules out the most common underlying cause
  5. 05If using a website platform or agency-managed site, ask explicitly whether any 'security' or 'bot blocking' feature has been enabled that affects these crawlers, since this is often switched on by default without discussion

Why Google can find a site but ChatGPT cannot

This is one of the most common support questions, and it usually comes down to one of three causes: OAI-SearchBot specifically blocked in robots.txt while Googlebot is allowed; a bot-protection or firewall service blocking OpenAI's user agents by default while allowing Google's (which many services treat as 'safe' by reputation); or content that only renders after client-side JavaScript runs, which Google processes (in a secondary rendering pass) more reliably than some AI crawlers are reported to. None of these is a penalty — they are configuration differences that are straightforward to check and fix.

Can AI crawlers read JavaScript-heavy websites?

Google documents that it renders JavaScript, but processes it in stages, and recommends not relying solely on client-side rendering for important content. It is widely reported — though not something every provider documents in full technical detail — that several AI-related crawlers do not reliably execute JavaScript, meaning content that only appears after a script runs may be invisible to them even though a human visitor sees it fine. The safest practice is ensuring core facts (what the company does, for whom, key services, contact details) exist in the server-rendered HTML, not only in a client-side-rendered component.

What is llms.txt, and do you need one?

llms.txt is a community-proposed file, similar in spirit to robots.txt or a sitemap, intended to give AI systems a concise summary of a site's key pages. It is not an official standard from any AI provider. As of current public information, no major AI search platform has publicly committed to reading or using it, and Google has stated it does not use it for Search. Creating one is low-cost and causes no harm, but there is no publicly documented evidence it improves AI visibility. Treat it as a minor, optional addition — not a priority ahead of crawlability, clear content and indexing fundamentals.

A simple decision framework

  • Visibility is the goal: allow GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Google-Extended, Googlebot and Bingbot unless there is a specific reason not to
  • Concerned specifically about model training on proprietary content: consider blocking GPTBot and Google-Extended only, leaving the search-facing crawlers (OAI-SearchBot, PerplexityBot, Googlebot, Bingbot) allowed
  • Not sure what is currently configured: check robots.txt, bot-protection settings and server logs before assuming either way
  • Site is JavaScript-heavy: confirm core facts exist in server-rendered HTML regardless of the robots.txt decision

Mistakes to avoid

  • Copying a blanket 'block all AI bots' robots.txt template without checking which named crawlers it actually lists
  • Assuming a CDN or security plugin's default settings allow these crawlers without checking
  • Treating llms.txt as a substitute for being properly indexed by Google and Bing
  • Reviewing robots.txt once and never again — platforms add and rename crawlers, and the right configuration is not fixed forever

Not sure why the website isn't being found or isn't producing enquiries?

A Commercial Website Audit reviews search visibility, structure, AI readiness and conversion together, and gives you a prioritised 90-day plan. £595 + VAT, fixed fee.

Related services

Written by

By Tom Evans

Founder, Evans Sales Consultancy

Published 4 October 2026 — 5 min read

Common questions

  • Not necessarily. GPTBot governs training data. OAI-SearchBot is the crawler that governs ChatGPT's search feature, per OpenAI's documentation — that is the one that matters for appearing in search-style ChatGPT answers.

  • Google states Google-Extended controls use for Gemini training and grounding only, and does not affect Google Search indexing or ranking.

  • Check server or CDN access logs for the named user agents, and review any bot-protection or WAF settings directly, since robots.txt alone does not show network-level blocking.

  • It is low-cost and not harmful, but no major AI platform has publicly committed to using it, and Google has said it does not use it for Search. It should not be prioritised over indexing and content fundamentals.

  • Some content on JavaScript-rendered sites may not be reliably read by AI crawlers, based on widely reported observations. Server-rendered HTML for key facts is the safer approach.

Still working out the right approach?

If your question is specific to your company, product or target market, we can help you work through the commercial options.

Discuss your market entry

More opportunities. Better conversion. Stronger sales. More revenue.

If your business could sell more than it currently does, the fastest way to find out why is to look at the numbers together.