Guide
robots.txt for AI bots: crawler rules and access checks
Review GPTBot, OAI-SearchBot and PerplexityBot rules, distinguish user-requested visits, and check whether a public page is actually reachable.
Last reviewed 30 September 2026
robots.txt states crawler preferences for a host. A public page also needs to be reachable through your server and CDN. Start with the intended search and training policies, then verify the served rules and page response. Check the public page for observable access barriers.
The agents worth naming
These are not interchangeable, and the distinction that catches people out is between the crawler that gathers pages in bulk and the fetcher that retrieves one page because a user asked a question right now.
| Token | Operator | What it does |
|---|---|---|
GPTBot | OpenAI | Content that may be used for model training |
OAI-SearchBot | OpenAI | Search index behind ChatGPT search |
ChatGPT-User | OpenAI | User-requested visits; robots.txt may not apply |
ClaudeBot | Anthropic | Bulk crawl |
Claude-Web, anthropic-ai | Anthropic | Older tokens, still seen |
PerplexityBot | Perplexity | Search discovery; not a foundation-model training crawler |
Perplexity-User | Perplexity | User-requested visits; generally ignores robots.txt |
Google-Extended | Controls Gemini use — not Google Search indexing | |
Applebot-Extended | Apple | Apple Intelligence |
CCBot | Common Crawl | Corpus many models train from |
OpenAI says robots.txt may not apply to user-initiated ChatGPT-User visits. Use OAI-SearchBot for the search crawl policy and GPTBot for the training crawl policy. Require authentication for private content. The OpenAI crawler comparison explains the separate choices.
How should I configure PerplexityBot?
Perplexity documents PerplexityBot for search discovery and Perplexity-User for user-requested visits. Review those separately. If your CDN allows a verified bot, use the provider's published IP information as well as the user agent; changing robots.txt alone does not remove a WAF challenge.
Longest match wins, and Allow breaks the tie
Google and Bing resolve rules by specificity, not by order: the longest matching pattern wins, and an Allow of equal length beats a Disallow. This is what makes it possible to open a single path inside a blocked directory.
User-agent: GPTBot Allow: /api/mcp Allow: /.well-known/ Disallow: /api/
If a public discovery document points to /api/mcp, a blanket /api/ disallow may prevent a compliant crawler from following it. Review the intended public paths before changing rules. The discovery guide explains the documents our scan inspects; it does not prove live MCP tool calls.
A file worth having
# Search User-agent: Googlebot Allow: / Disallow: /admin/ User-agent: Bingbot Allow: / Disallow: /admin/ # AI crawlers and live fetchers User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: PerplexityBot User-agent: Google-Extended Allow: / Allow: /.well-known/ Disallow: /admin/ Disallow: /cart/ Disallow: /account/ # Everything else User-agent: * Allow: / Disallow: /admin/ Sitemap: https://example.com/sitemap.xml
Consecutive User-agent lines form one group, which keeps the file readable. Do not disallow /api/ wholesale if anything under it is meant to be found.
Why we mark the catch-all as partial
A site whose only rule is User-agent: * with Allow: / does permit every AI crawler. Our check reports that as partial rather than a pass, and the reason is not pedantry: it means nobody has made a decision. The next person to add a blanket Disallow for an unrelated reason will silently take out the assistants too, and nothing in your analytics will show it.
Naming the agents is a small act of documentation. It says we considered this, and it makes the next change deliberate.
robots.txt is a request, not a control
Crawler preferences and server access controls need separate checks. The scanner sends agent user-agent strings and observes responses; it does not prove that those requests originate from the named provider. Compare the finding with your CDN rules and legitimate-request logs.
If you want to state usage terms rather than block outright, content signals in robots.txt let you express intent — search yes, training no, for instance — without removing yourself from the answers.
Sources
Source references · article reviewed 30 September 2026
Keep reading
Check your own site against this
The Agent Readiness Score measures exactly what this article describes, and shows the evidence behind every finding.
Run the check →