Guides

robots.txt for AI crawlers, one bot at a time

AI crawlers do three different jobs, and each vendor names a separate token for each. A table built from the vendors' own pages, the RFC rules that trip people up, and three files to copy.

By The Laarpi teamUpdated 7 min read

AI crawlers do three different jobs: collecting training data, building a search index for an assistant, and fetching a page because a person just asked about it. OpenAI, Anthropic and Perplexity each give those jobs separate user-agent tokens, so you can admit one and refuse another. Write one robots.txt group per token you want to treat differently, repeat your private paths inside each group, and then check that your CDN agrees with the file. Admitting the right crawlers is one part of an AI-ready website; this guide covers that one file, with every fact taken from the vendor's own page as of 6 October 2026.

Three jobs, three kinds of token

  • Training crawlers collect pages to train future models. Blocking them keeps future content out of training sets. It doesn't remove you from anything people see today.
  • Search crawlers build the index an assistant searches when it answers with links. Blocking them keeps you out of those answers.
  • User-triggered fetchers visit one page because a person pasted a link or asked a question. Some vendors say robots.txt may not apply to these, because a human made the request.

Two tokens in the list below aren't crawlers at all. Google-Extended and Applebot-Extended are control tokens: Google's and Apple's normal crawlers do the fetching, and the token only tells them what the content may be used for.

Every AI crawler, from its owner's page

TokenOwnerJobWhat blocking it does, in the owner's wordsObeys robots.txt
GPTBotOpenAITraining"indicates a site's content should not be used in training"Yes
OAI-SearchBotOpenAISearchBlocked sites "will not be shown in ChatGPT search answers"Yes
ChatGPT-UserOpenAIUser fetchNot stated"robots.txt rules may not apply"
ClaudeBotAnthropicTrainingFuture materials "excluded from our AI model training datasets"Yes
Claude-SearchBotAnthropicSearch"prevents our system from indexing your content for search optimization"Yes
Claude-UserAnthropicUser fetch"prevents our system from retrieving your content in response to a user query"Yes
PerplexityBotPerplexitySearchNot shown in Perplexity's search results; "not used to crawl content for AI foundation models"Yes
Perplexity-UserPerplexityUser fetchNot stated"generally ignores robots.txt rules"
Google-ExtendedGoogleControl tokenContent not used for training Gemini models or for grounding; "does not impact a site's inclusion in Google Search"Read by Google's crawlers
Applebot-ExtendedAppleControl tokenContent not used to train Apple's foundation models; "does not crawl webpages"Read by Applebot
meta-externalagentMetaTraining and indexingCrawls "for use cases such as training foundation AI models or improving products by indexing content directly"Yes
CCBotCommon CrawlOpen datasetLeaves you out of Common Crawl's public crawl archiveYes

Three details from the same pages are worth knowing:

  • Anthropic says all three of its bots honour robots.txt, including the user fetcher, and it supports the non-standard Crawl-delay extension.
  • Meta's user fetcher, meta-externalfetcher, "may bypass robots.txt rules", like ChatGPT-User and Perplexity-User.
  • Applebot follows your Googlebot rules when you don't have an Applebot group. Apple says it "will follow Googlebot instructions."

Where Google's AI answers fit

Google-Extended is the token people reach for, and it is often the wrong one. Google says it controls whether crawled content may be used to train Gemini models and for grounding, and that it doesn't affect inclusion or ranking in Google Search. Google's page on AI features says a page appears in them if it is "indexed and eligible to be shown in Google Search with a snippet", and points to nosnippet, data-nosnippet, max-snippet and noindex as the controls. Refusing Google-Extended doesn't remove you from AI Overviews. Refusing Googlebot removes you from Google Search entirely.

The RFC 9309 rules that catch people out

robots.txt became an IETF standard, RFC 9309, in September 2022. Most mistakes come from six of its rules.

  1. A named group replaces the * group. A crawler obeys the group that names it and ignores *. If you add User-agent: GPTBot with only Allow: /, GPTBot may now crawl the paths you disallowed for everyone else.
  2. Groups for the same crawler are merged. If GPTBot appears in two groups, "the matching groups' rules MUST be combined into one group".
  3. Token matching ignores case. gptbot and GPTBot are the same group.
  4. The longest matching rule wins. "The most specific match is the match that has the most octets." When an Allow and a Disallow match equally, the Allow should win.
  5. The status code of the file itself matters. If /robots.txt returns a 4xx, a crawler "MAY access any resources on the server". If it returns a 5xx, the crawler "MUST assume complete disallow". A broken robots route that throws a server error blocks every crawler.
  6. Changes aren't instant, and size has a floor. Crawlers should not use a cached copy for more than 24 hours, and they must parse at least 500 KiB.

Here is rule 1 going wrong, then right:

# Wrong: GPTBot ignores the * group, so /account/ is open to it
User-agent: *
Disallow: /account/

User-agent: GPTBot
Allow: /

# Right: the named group repeats the private path
User-agent: *
Disallow: /account/

User-agent: GPTBot
Allow: /
Disallow: /account/

Three files to copy

Replace /account/ with your own private paths and example.com with your domain. One group can list several User-agent lines.

1. Allow every AI use. You don't need named groups for this at all: the * group covers every crawler without one.

User-agent: *
Allow: /
Disallow: /account/

Sitemap: https://example.com/sitemap.xml

2. Allow AI search and user fetches, refuse training. This keeps you in assistants' search answers and lets them read a page a person asks about, while asking the training crawlers to stay out.

User-agent: *
Allow: /
Disallow: /account/

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /

Sitemap: https://example.com/sitemap.xml

Weigh the last row on its own: Common Crawl's archive is an open repository that anyone can use, for research as much as for model training, so refusing CCBot leaves you out of all of it.

3. Refuse every AI crawler, keep classic search. Googlebot, Bingbot and Applebot keep their access through the * group, so your search listings stay.

User-agent: *
Allow: /
Disallow: /account/

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /

Sitemap: https://example.com/sitemap.xml

File 3 still can't bind the fetchers that say robots.txt may not apply to them. If you need those out too, block them at the edge.

Your CDN can overrule the file

Since 15 September 2026, Cloudflare offers new domains one of two presets. A site that earns money from ads gets "Disallow AI Training", with agents blocked on pages that serve an ad. A site without ads gets "Allow" for both. Existing domains keep the settings they had. Cloudflare's Bot Preference Sync, announced on 21 August 2026 and on by default for new customers, writes those choices into your robots.txt, and its managed robots.txt adds its lines before your own file's content.

Two consequences follow. The robots.txt in your repository may not be the one crawlers receive, so always check the served file. And a crawler your file admits can still be stopped by a firewall rule or a challenge page. Why AI can't see your website walks through that layer.

Preference signals beyond allow and disallow

Two newer layers say what content may be used for, not who may fetch it. Neither is access control.

  • Cloudflare's Content Signals add a line to robots.txt with three terms. search covers building a search index and showing results. ai-input covers putting content into AI models in real time, as in grounding and AI search answers. ai-train covers training or fine-tuning models. A line looks like Content-Signal: search=yes, ai-input=yes, ai-train=no.
  • The IETF's AIPREF working group is drafting a shared vocabulary. The current revision, draft-ietf-aipref-vocab-08 (14 September 2026), is an Internet-Draft, which by its own text is work in progress. The draft also says it doesn't "ensure that preferences are followed".

Check what you actually serve

# 1. The file crawlers get, and its status code (expect 200)
curl -s -o /dev/null -w "%{http_code}\n" https://example.com/robots.txt
curl -s https://example.com/robots.txt

# 2. What your edge does with a crawler's user agent
curl -s -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
  https://example.com/

A forged user agent only tests rules that match on the user agent. A CDN that verifies real crawlers by IP will treat your request differently from GPTBot's. OpenAI publishes the address ranges at openai.com/gptbot.json, openai.com/searchbot.json and openai.com/chatgpt-user.json, and Perplexity at perplexity.com/perplexitybot.json.

The free AI-ready check reads your served robots.txt the way RFC 9309 describes, including group merging and longest match, and reports whether GPTBot, ClaudeBot, PerplexityBot and Google-Extended may reach your public pages.

A worked example: this site's own file

Laarpi's own robots.txt allows everything public and keeps the studio and the API private. It names 16 AI tokens in one group, because a named group replaces the * group, and the comment in the source says exactly that:

User-Agent: *
Allow: /
Disallow: /studio
Disallow: /api/

User-Agent: GPTBot
User-Agent: OAI-SearchBot
User-Agent: ChatGPT-User
User-Agent: ClaudeBot
User-Agent: Claude-User
User-Agent: Claude-SearchBot
User-Agent: anthropic-ai
User-Agent: PerplexityBot
User-Agent: Perplexity-User
User-Agent: Google-Extended
User-Agent: Applebot-Extended
User-Agent: Amazonbot
User-Agent: CCBot
User-Agent: meta-externalagent
User-Agent: DuckAssistBot
User-Agent: MistralAI-User
Allow: /
Disallow: /studio
Disallow: /api/

Listing tokens that only get what * already gives them changes nothing for those crawlers. It does make the policy explicit, and it means a later edit to one group can't quietly open the private paths. robots.txt is only the first thing a crawler reads. For what it does once it's in, read how AI agents read websites, and for the rest of the agent-facing layer (llms.txt, Markdown copies of each page, structured data and agent tools) see what makes a website AI-ready.

Sources

Checked 6 October 2026. Standards and vendor pages change; the linked pages are the authority.

  1. RFC 9309: Robots Exclusion Protocol (IETF, September 2022)
  2. OpenAI: overview of OpenAI crawlers
  3. Anthropic: does Anthropic crawl data from the web? (updated 7 April 2026)
  4. Perplexity: Perplexity crawlers
  5. Google Search Central: Google's common crawlers, Google-Extended (updated 14 July 2026)
  6. Google Search Central: AI features and your website (updated 10 December 2025)
  7. Apple: about Applebot (updated 4 September 2026)
  8. Meta: Meta web crawlers
  9. Common Crawl: CCBot
  10. Cloudflare: stay discoverable in search while disallowing AI training (15 September 2026)
  11. Cloudflare: introducing Bot Preference Sync (21 August 2026)
  12. Cloudflare docs: managed robots.txt and Content Signals (updated 3 August 2026)
  13. IETF: draft-ietf-aipref-vocab-08 (Internet-Draft, 14 September 2026)
Questions

Fair questions

If I block GPTBot, will my site disappear from ChatGPT search?

No. OpenAI documents two separate crawlers. GPTBot collects training data, and blocking it 'indicates a site's content should not be used in training'. OAI-SearchBot surfaces sites in ChatGPT search, and only blocking that one keeps you out of ChatGPT's search answers.

Does blocking Google-Extended remove my site from Google Search or AI Overviews?

Google says Google-Extended 'does not impact a site's inclusion in Google Search nor is it used as a ranking signal'. It controls use for training Gemini models and for grounding. For Google's AI features in Search, Google points to the usual snippet controls (nosnippet, data-nosnippet, max-snippet) and noindex.

Is robots.txt enough to keep AI bots out?

No. RFC 9309 says the rules 'are not a form of access authorization'. OpenAI says robots.txt 'may not apply' to ChatGPT-User, Perplexity says Perplexity-User 'generally ignores' it, and Meta says meta-externalfetcher 'may bypass' it, because a person asked for the page. To keep a page private, put it behind sign-in or block it at your server or CDN.

How long does a robots.txt change take to work?

RFC 9309 says crawlers should not use a cached copy for more than 24 hours. OpenAI says it can take about 24 hours for its systems to adjust, and Perplexity says up to 24 hours.

Do I need a group for every AI crawler?

Only for the ones you want to treat differently from everyone else. A crawler with no group of its own follows the * group. If you do add a named group, it replaces the * group for that crawler, so repeat your private Disallow lines inside it.

What is the difference between ClaudeBot, Claude-SearchBot and Claude-User?

Anthropic documents ClaudeBot as the crawler that collects training data, Claude-SearchBot as the one that indexes for search results, and Claude-User as the fetcher that visits a page when someone asks Claude about it. Anthropic says all its bots honour robots.txt, so you can treat each one separately.

Related

Start building

Describe the site in a sentence. It asks what matters, then designs and builds it from scratch.

One sentence is enough.