robots.txt for AI crawlers, one bot at a time
AI crawlers do three different jobs, and each vendor names a separate token for each. A table built from the vendors' own pages, the RFC rules that trip people up, and three files to copy.
AI crawlers do three different jobs: collecting training data, building a search index for an assistant, and fetching a page because a person just asked about it. OpenAI, Anthropic and Perplexity each give those jobs separate user-agent tokens, so you can admit one and refuse another. Write one robots.txt group per token you want to treat differently, repeat your private paths inside each group, and then check that your CDN agrees with the file. Admitting the right crawlers is one part of an AI-ready website; this guide covers that one file, with every fact taken from the vendor's own page as of 6 October 2026.
Three jobs, three kinds of token
- Training crawlers collect pages to train future models. Blocking them keeps future content out of training sets. It doesn't remove you from anything people see today.
- Search crawlers build the index an assistant searches when it answers with links. Blocking them keeps you out of those answers.
- User-triggered fetchers visit one page because a person pasted a link or asked a question. Some vendors say robots.txt may not apply to these, because a human made the request.
Two tokens in the list below aren't crawlers at all. Google-Extended and Applebot-Extended are control tokens: Google's and Apple's normal crawlers do the fetching, and the token only tells them what the content may be used for.
Every AI crawler, from its owner's page
| Token | Owner | Job | What blocking it does, in the owner's words | Obeys robots.txt |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | "indicates a site's content should not be used in training" | Yes |
| OAI-SearchBot | OpenAI | Search | Blocked sites "will not be shown in ChatGPT search answers" | Yes |
| ChatGPT-User | OpenAI | User fetch | Not stated | "robots.txt rules may not apply" |
| ClaudeBot | Anthropic | Training | Future materials "excluded from our AI model training datasets" | Yes |
| Claude-SearchBot | Anthropic | Search | "prevents our system from indexing your content for search optimization" | Yes |
| Claude-User | Anthropic | User fetch | "prevents our system from retrieving your content in response to a user query" | Yes |
| PerplexityBot | Perplexity | Search | Not shown in Perplexity's search results; "not used to crawl content for AI foundation models" | Yes |
| Perplexity-User | Perplexity | User fetch | Not stated | "generally ignores robots.txt rules" |
| Google-Extended | Control token | Content not used for training Gemini models or for grounding; "does not impact a site's inclusion in Google Search" | Read by Google's crawlers | |
| Applebot-Extended | Apple | Control token | Content not used to train Apple's foundation models; "does not crawl webpages" | Read by Applebot |
| meta-externalagent | Meta | Training and indexing | Crawls "for use cases such as training foundation AI models or improving products by indexing content directly" | Yes |
| CCBot | Common Crawl | Open dataset | Leaves you out of Common Crawl's public crawl archive | Yes |
Three details from the same pages are worth knowing:
- Anthropic says all three of its bots honour robots.txt, including the user fetcher, and it supports the non-standard
Crawl-delayextension. - Meta's user fetcher, meta-externalfetcher, "may bypass robots.txt rules", like ChatGPT-User and Perplexity-User.
- Applebot follows your Googlebot rules when you don't have an Applebot group. Apple says it "will follow Googlebot instructions."
Where Google's AI answers fit
Google-Extended is the token people reach for, and it is often the wrong one. Google says it controls whether crawled content may be used to train Gemini models and for grounding, and that it doesn't affect inclusion or ranking in Google Search. Google's page on AI features says a page appears in them if it is "indexed and eligible to be shown in Google Search with a snippet", and points to nosnippet, data-nosnippet, max-snippet and noindex as the controls. Refusing Google-Extended doesn't remove you from AI Overviews. Refusing Googlebot removes you from Google Search entirely.
The RFC 9309 rules that catch people out
robots.txt became an IETF standard, RFC 9309, in September 2022. Most mistakes come from six of its rules.
- A named group replaces the
*group. A crawler obeys the group that names it and ignores*. If you addUser-agent: GPTBotwith onlyAllow: /, GPTBot may now crawl the paths you disallowed for everyone else. - Groups for the same crawler are merged. If GPTBot appears in two groups, "the matching groups' rules MUST be combined into one group".
- Token matching ignores case.
gptbotandGPTBotare the same group. - The longest matching rule wins. "The most specific match is the match that has the most octets." When an Allow and a Disallow match equally, the Allow should win.
- The status code of the file itself matters. If
/robots.txtreturns a 4xx, a crawler "MAY access any resources on the server". If it returns a 5xx, the crawler "MUST assume complete disallow". A broken robots route that throws a server error blocks every crawler. - Changes aren't instant, and size has a floor. Crawlers should not use a cached copy for more than 24 hours, and they must parse at least 500 KiB.
Here is rule 1 going wrong, then right:
# Wrong: GPTBot ignores the * group, so /account/ is open to it
User-agent: *
Disallow: /account/
User-agent: GPTBot
Allow: /
# Right: the named group repeats the private path
User-agent: *
Disallow: /account/
User-agent: GPTBot
Allow: /
Disallow: /account/
Three files to copy
Replace /account/ with your own private paths and example.com with your domain. One group can list several User-agent lines.
1. Allow every AI use. You don't need named groups for this at all: the * group covers every crawler without one.
User-agent: *
Allow: /
Disallow: /account/
Sitemap: https://example.com/sitemap.xml
2. Allow AI search and user fetches, refuse training. This keeps you in assistants' search answers and lets them read a page a person asks about, while asking the training crawlers to stay out.
User-agent: *
Allow: /
Disallow: /account/
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /
Sitemap: https://example.com/sitemap.xml
Weigh the last row on its own: Common Crawl's archive is an open repository that anyone can use, for research as much as for model training, so refusing CCBot leaves you out of all of it.
3. Refuse every AI crawler, keep classic search. Googlebot, Bingbot and Applebot keep their access through the * group, so your search listings stay.
User-agent: *
Allow: /
Disallow: /account/
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
User-agent: CCBot
Disallow: /
Sitemap: https://example.com/sitemap.xml
File 3 still can't bind the fetchers that say robots.txt may not apply to them. If you need those out too, block them at the edge.
Your CDN can overrule the file
Since 15 September 2026, Cloudflare offers new domains one of two presets. A site that earns money from ads gets "Disallow AI Training", with agents blocked on pages that serve an ad. A site without ads gets "Allow" for both. Existing domains keep the settings they had. Cloudflare's Bot Preference Sync, announced on 21 August 2026 and on by default for new customers, writes those choices into your robots.txt, and its managed robots.txt adds its lines before your own file's content.
Two consequences follow. The robots.txt in your repository may not be the one crawlers receive, so always check the served file. And a crawler your file admits can still be stopped by a firewall rule or a challenge page. Why AI can't see your website walks through that layer.
Preference signals beyond allow and disallow
Two newer layers say what content may be used for, not who may fetch it. Neither is access control.
- Cloudflare's Content Signals add a line to robots.txt with three terms.
searchcovers building a search index and showing results.ai-inputcovers putting content into AI models in real time, as in grounding and AI search answers.ai-traincovers training or fine-tuning models. A line looks likeContent-Signal: search=yes, ai-input=yes, ai-train=no. - The IETF's AIPREF working group is drafting a shared vocabulary. The current revision, draft-ietf-aipref-vocab-08 (14 September 2026), is an Internet-Draft, which by its own text is work in progress. The draft also says it doesn't "ensure that preferences are followed".
Check what you actually serve
# 1. The file crawlers get, and its status code (expect 200)
curl -s -o /dev/null -w "%{http_code}\n" https://example.com/robots.txt
curl -s https://example.com/robots.txt
# 2. What your edge does with a crawler's user agent
curl -s -o /dev/null -w "%{http_code}\n" \
-A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
https://example.com/
A forged user agent only tests rules that match on the user agent. A CDN that verifies real crawlers by IP will treat your request differently from GPTBot's. OpenAI publishes the address ranges at openai.com/gptbot.json, openai.com/searchbot.json and openai.com/chatgpt-user.json, and Perplexity at perplexity.com/perplexitybot.json.
The free AI-ready check reads your served robots.txt the way RFC 9309 describes, including group merging and longest match, and reports whether GPTBot, ClaudeBot, PerplexityBot and Google-Extended may reach your public pages.
A worked example: this site's own file
Laarpi's own robots.txt allows everything public and keeps the studio and the API private. It names 16 AI tokens in one group, because a named group replaces the * group, and the comment in the source says exactly that:
User-Agent: *
Allow: /
Disallow: /studio
Disallow: /api/
User-Agent: GPTBot
User-Agent: OAI-SearchBot
User-Agent: ChatGPT-User
User-Agent: ClaudeBot
User-Agent: Claude-User
User-Agent: Claude-SearchBot
User-Agent: anthropic-ai
User-Agent: PerplexityBot
User-Agent: Perplexity-User
User-Agent: Google-Extended
User-Agent: Applebot-Extended
User-Agent: Amazonbot
User-Agent: CCBot
User-Agent: meta-externalagent
User-Agent: DuckAssistBot
User-Agent: MistralAI-User
Allow: /
Disallow: /studio
Disallow: /api/
Listing tokens that only get what * already gives them changes nothing for those crawlers. It does make the policy explicit, and it means a later edit to one group can't quietly open the private paths. robots.txt is only the first thing a crawler reads. For what it does once it's in, read how AI agents read websites, and for the rest of the agent-facing layer (llms.txt, Markdown copies of each page, structured data and agent tools) see what makes a website AI-ready.
Sources
Checked 6 October 2026. Standards and vendor pages change; the linked pages are the authority.
- RFC 9309: Robots Exclusion Protocol (IETF, September 2022)
- OpenAI: overview of OpenAI crawlers
- Anthropic: does Anthropic crawl data from the web? (updated 7 April 2026)
- Perplexity: Perplexity crawlers
- Google Search Central: Google's common crawlers, Google-Extended (updated 14 July 2026)
- Google Search Central: AI features and your website (updated 10 December 2025)
- Apple: about Applebot (updated 4 September 2026)
- Meta: Meta web crawlers
- Common Crawl: CCBot
- Cloudflare: stay discoverable in search while disallowing AI training (15 September 2026)
- Cloudflare: introducing Bot Preference Sync (21 August 2026)
- Cloudflare docs: managed robots.txt and Content Signals (updated 3 August 2026)
- IETF: draft-ietf-aipref-vocab-08 (Internet-Draft, 14 September 2026)
