Guide
GPTBot vs OAI-SearchBot vs ChatGPT-User: robots.txt rules
Choose crawler rules for training, ChatGPT search and user-requested visits separately. Inspect the public response and verify the edge configuration.
Last reviewed 30 September 2026
GPTBot, OAI-SearchBot and ChatGPT-User have different purposes. A rule for one is not a policy for all three. Start by deciding which public pages you want discovered, which training policy you intend, and which access restrictions your server must enforce. Check a public page for observed access barriers before editing your rules.
Which OpenAI user agent does what?
| User agent | Documented purpose | What to decide |
|---|---|---|
GPTBot | Crawls content that may be used for model training | Your training crawl policy |
OAI-SearchBot | Discovers content for ChatGPT search | Your search crawl policy |
ChatGPT-User | Visits associated with user actions | Real access controls; robots.txt may not apply |
Can I allow ChatGPT search while opting out of GPTBot?
OpenAI documents independent settings for its search and training crawlers. The following example permits the search crawler and opts out of the training crawler. It is a policy example for public content, not an access-control configuration or a guarantee of appearing in answers.
# Example: public search discovery allowed, training crawl opted out. User-agent: OAI-SearchBot Allow: / User-agent: GPTBot Disallow: /
Preserve existing path restrictions, sitemap declarations and rules for other crawlers. Inspect matching groups before adding another one. The robots.txt guide covers path matching and the limits of a catch-all rule.
Does blocking ChatGPT-User prevent user-requested visits?
OpenAI says robots.txt rules may not apply to user-initiated ChatGPT-User requests. It is not the control for search inclusion. Require authentication and enforce permissions for private data; do not use a robots.txt disallow as the security boundary.
How do I verify the change?
- Fetch the deployed
/robots.txtfrom the exact hostname. Confirm its status, text and the intended groups. - Inspect the target page response and your CDN or WAF rules. A page can permit a crawler in robots.txt while still returning a challenge or 403.
- For an allow rule based on bot identity, check the provider's current published IP information. A user-agent string alone can be spoofed.
- Repeat the public scan, then record the specific finding that changed. Track actual search referrals separately; a passed access check does not prove a citation.
What can the checker tell me?
The WebMCP checker reports public source and access signals. It cannot establish the identity of every requester, prove inclusion in a search index or measure a future mention. Read the scoring method before interpreting a crawler finding.
Sources
Source references · article reviewed 30 September 2026
- OpenAI — Overview of crawlers — Search, training, user-requested visits and published IP references
- RFC 9309 — Robots Exclusion Protocol — Group and path matching; access-control limitations
Keep reading
Check your own site against this
The Agent Readiness Score measures exactly what this article describes, and shows the evidence behind every finding.
Run the check →