webmcp-tool

Guide

GPTBot vs OAI-SearchBot vs ChatGPT-User: robots.txt rules

Choose crawler rules for training, ChatGPT search and user-requested visits separately. Inspect the public response and verify the edge configuration.

Last reviewed 30 September 2026

GPTBot, OAI-SearchBot and ChatGPT-User have different purposes. A rule for one is not a policy for all three. Start by deciding which public pages you want discovered, which training policy you intend, and which access restrictions your server must enforce. Check a public page for observed access barriers before editing your rules.

Which OpenAI user agent does what?

User agentDocumented purposeWhat to decide
GPTBotCrawls content that may be used for model trainingYour training crawl policy
OAI-SearchBotDiscovers content for ChatGPT searchYour search crawl policy
ChatGPT-UserVisits associated with user actionsReal access controls; robots.txt may not apply

Can I allow ChatGPT search while opting out of GPTBot?

OpenAI documents independent settings for its search and training crawlers. The following example permits the search crawler and opts out of the training crawler. It is a policy example for public content, not an access-control configuration or a guarantee of appearing in answers.

# Example: public search discovery allowed, training crawl opted out.
User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /
Merge these groups into your existing robots.txt after reviewing its policy.

Preserve existing path restrictions, sitemap declarations and rules for other crawlers. Inspect matching groups before adding another one. The robots.txt guide covers path matching and the limits of a catch-all rule.

Does blocking ChatGPT-User prevent user-requested visits?

OpenAI says robots.txt rules may not apply to user-initiated ChatGPT-User requests. It is not the control for search inclusion. Require authentication and enforce permissions for private data; do not use a robots.txt disallow as the security boundary.

How do I verify the change?

  1. Fetch the deployed /robots.txt from the exact hostname. Confirm its status, text and the intended groups.
  2. Inspect the target page response and your CDN or WAF rules. A page can permit a crawler in robots.txt while still returning a challenge or 403.
  3. For an allow rule based on bot identity, check the provider's current published IP information. A user-agent string alone can be spoofed.
  4. Repeat the public scan, then record the specific finding that changed. Track actual search referrals separately; a passed access check does not prove a citation.

What can the checker tell me?

The WebMCP checker reports public source and access signals. It cannot establish the identity of every requester, prove inclusion in a search index or measure a future mention. Read the scoring method before interpreting a crawler finding.

Sources

Source references · article reviewed 30 September 2026

  1. OpenAI — Overview of crawlers — Search, training, user-requested visits and published IP references
  2. RFC 9309 — Robots Exclusion Protocol — Group and path matching; access-control limitations

Keep reading

Check your own site against this

The Agent Readiness Score measures exactly what this article describes, and shows the evidence behind every finding.

Run the check →