Quick answer
AI companies run different crawlers for training, for building a search index and for fetching a page when a user asks. To appear in AI search answers, do not block the search crawlers. Whether to allow training crawlers is a business decision. Always confirm current names in each vendor's documentation.
Key takeaways
- Separate three purposes: training, search indexing and user-triggered fetching.
- Blocking a search crawler can remove you from that assistant's search answers.
- Bot names change. Check the vendor documentation before editing robots.txt.
What kinds of AI crawlers exist?
| Vendor | Example user agent | Documented purpose |
|---|---|---|
| OpenAI | OAI-SearchBot | Surfacing sites in ChatGPT search features |
| OpenAI | GPTBot | Collecting content that may be used to train models |
| OpenAI | ChatGPT-User | Fetching a page when a user asks ChatGPT to |
| Perplexity | PerplexityBot | Building the Perplexity search index |
| Perplexity | Perplexity-User | Fetching a page when a user asks |
| Google-Extended | A control token for whether content may be used for some Google AI models. Ordinary Googlebot handles Search. |
How do you decide what to allow?
- 1List every bot in your server logs that identifies as an AI crawler.
- 2For each, read the vendor's page and record its stated purpose.
- 3Allow crawlers that build search indexes for platforms where you want to be visible.
- 4Decide separately about training crawlers, based on your content and legal position.
- 5Update robots.txt, then check your logs to confirm the change took effect.
An example robots.txt fragment
- User-agent: OAI-SearchBot followed by Allow: / lets that search crawler read the site.
- User-agent: PerplexityBot followed by Allow: / does the same for Perplexity.
- User-agent: GPTBot followed by Disallow: / opts out of that training crawler only.
robots.txt is a request, not a lock
The documented agents in one place
| Provider | Agent | Documented purpose | robots.txt |
|---|---|---|---|
| OpenAI | OAI-SearchBot | Surface websites in ChatGPT's search features | Disallowing keeps a site out of ChatGPT search answers, though it can still appear as a navigational link; changes take about 24 hours |
| OpenAI | GPTBot | Crawl content that may be used to train foundation models | Disallowing indicates content should not be used for training |
| OpenAI | ChatGPT-User | Certain user actions in ChatGPT and Custom GPTs | User-initiated, so robots.txt rules may not apply |
| Perplexity | PerplexityBot | Surface and link websites in Perplexity results; not used to train foundation models | Respected; allowing it is recommended to appear in results |
| Perplexity | Perplexity-User | Visit a page when a user asks | Generally ignores robots.txt because the fetch is user-requested |
| Anthropic | ClaudeBot | Collect content that may contribute to training | Restricting excludes future materials from training |
| Anthropic | Claude-User | Retrieve pages for a person's question | Disabling may reduce visibility for user-directed search |
| Anthropic | Claude-SearchBot | Improve search result quality | Disabling may reduce visibility in search results |
Google uses its own crawler for Search, and documents that robots.txt directives for Googlebot are the control for access. Names and behaviour change, so treat the provider pages listed in the sources as the authority.
Three decisions, not one
- 1Do you want to be eligible for search-connected answers? If yes, allow the search crawlers.
- 2Do you want your content used for model training? This is a business and legal decision, made separately.
- 3Do you need to keep certain content out of reach entirely? Use authentication, because robots.txt is a request and does not bind user-initiated fetches.
A worked robots.txt pattern
A site that wants to appear in search-connected answers but opt out of training could allow OAI-SearchBot and PerplexityBot and disallow GPTBot and ClaudeBot, then leave the rules for user-initiated agents to the provider documentation. The exact rules should be written and tested against each provider's current documentation before deployment. This describes a pattern and is not advice for any particular site.
Verify in your logs
- Search your server or CDN logs for each user agent you have allowed or blocked.
- Confirm blocked agents stop after the documented delay.
- Confirm allowed agents can fetch your key pages and receive normal responses.
- Check that a CDN or firewall rule is not blocking crawlers you meant to allow.
A policy your team can write down
Turn the three decisions into a short written policy: which search crawlers are allowed, what the position on training crawlers is and why, and how private content is protected. Give it an owner and a review date. Writing it down prevents the common situation where a robots.txt file accumulates rules nobody remembers adding.
Review the policy whenever a provider changes its documentation, and at least once a quarter. Crawler names, purposes and behaviour are not fixed, and a rule that made sense last year may block something you now want, or allow something you now would not.
Will allowing a crawler guarantee my brand is cited?
No. It only makes citation possible. Content quality, relevance and third party corroboration still decide whether you are named.
Can I block training but stay in AI search?
Vendors that publish separate crawlers for each purpose let you do this. Confirm the current behaviour in their documentation.
Where do I check this on my own site?
Open yourdomain/robots.txt and review your server logs for AI user agents.
How long after editing robots.txt does the change take effect?
It depends on the provider. OpenAI documents about 24 hours for OAI-SearchBot. Check each provider's documentation.
Can I block only some paths?
Yes. robots.txt supports path-specific rules. Keep them simple and test them.
Sources
- 1. OpenAI: Overview of OpenAI crawlers (Accessed Sep 2026)
- 2. Perplexity: Crawlers (Accessed Sep 2026)
- 3. Google Search Central: Introduction to robots.txt (Accessed Sep 2026)
- 4. Google Search Central: AI features and your website (Accessed Sep 2026)
- 5. OpenAI: Overview of OpenAI crawlers (Accessed Sep 2026)
- 6. Anthropic Help Center: Does Anthropic crawl data from the web, and how can site owners block the crawler? (Accessed Sep 2026)
Cite this page
BrandWater AI Research. (2026, 18 August 2026). AI crawlers and robots.txt: which bots to allow, and what each one does. https://brandwaterai.in/guides/ai-crawlers-and-robots-txt
How we work
Figures are dated and linked to their sources. Where none exist we say so. Read our methodology and AI transparency pages.