Quick answer
Major AI vendors document separate crawlers for training data, search indexing and fetching a page at a user's request. Because the purposes differ, blocking one does not necessarily block the others. Names and behaviour change, so confirm them in the vendor documentation linked here.
| Vendor | Search index | Training | User-triggered |
|---|---|---|---|
| OpenAI | OAI-SearchBot | GPTBot | ChatGPT-User |
| Perplexity | PerplexityBot | Not stated as a separate crawler | Perplexity-User |
| Googlebot (Search) | Google-Extended control token | Google-specific fetchers |
Method note
What does this mean for visibility?
A brand that blocks a vendor's search crawler may be left out of that vendor's search grounded answers. A brand that blocks a training crawler is making a separate decision. Treat the two independently.
Adding Anthropic and the user-initiated fetches
The table above lists OpenAI, Perplexity and Google. Anthropic documents three agents as well. ClaudeBot collects web content that may contribute to training. Claude-User supports Claude users: when a person asks a question, it may access websites. Claude-SearchBot navigates the web to improve search result quality. Anthropic explains that restricting ClaudeBot signals that future materials should be excluded from training datasets, while disabling Claude-User or Claude-SearchBot may reduce visibility in user-directed search and search results respectively.
| Purpose | OpenAI | Perplexity | Anthropic |
|---|---|---|---|
| Search indexing | OAI-SearchBot | PerplexityBot | Claude-SearchBot |
| Model training | GPTBot | Not used to crawl for foundation models, per Perplexity | ClaudeBot |
| User-initiated fetch | ChatGPT-User | Perplexity-User | Claude-User |
What the documentation says about robots.txt
- OpenAI: disallowing OAI-SearchBot keeps a site out of ChatGPT search answers, though it can still appear as a navigational link, and changes take about 24 hours to process. Each setting is independent of the others.
- OpenAI: because ChatGPT-User actions are initiated by a user, robots.txt rules may not apply.
- Perplexity: PerplexityBot respects robots.txt and is recommended if you want to appear in results. Perplexity-User generally ignores robots.txt because a user requested the fetch.
- Anthropic: each agent is a separate robots.txt token, and restricting one does not restrict the others.
Why this matters more than it looks
The most common mistake is treating all AI bots as one thing. A site that blocks everything to protect its content from training can also remove itself from search-connected answers. A site that allows everything may be making a decision about training use that it did not intend. Separating the three purposes lets each be decided on its own merits.
How to check your own setup
- 1Read your current robots.txt and list every AI-related user agent it mentions.
- 2Match each to the provider's documentation and write down its purpose.
- 3Decide the search, training and user-fetch settings separately for each provider.
- 4After changing rules, check server logs to confirm the behaviour changed after the documented delay.
- 5Review quarterly, because names and behaviour change.
robots.txt is a request
Where can I see which bots visit my site?
In your server or CDN logs. Filter by the user agent strings listed in each vendor's documentation.
Are these lists complete?
No. New crawlers appear, so treat the vendor pages as the source of truth.
Will allowing search crawlers expose my content to training?
The providers document search and training crawlers as separate agents with independent settings. Confirm each provider's current documentation, because policies can change.
What about crawlers I have never heard of?
Log them, check the provider's documentation, and treat unknown agents with caution. New crawlers appear regularly.
Sources
- 1. OpenAI: Overview of OpenAI crawlers (Accessed Sep 2026)
- 2. Perplexity: Crawlers (Accessed Sep 2026)
- 3. Google Search Central: AI features and your website (Accessed Sep 2026)
- 4. Google Search Central: Introduction to robots.txt (Accessed Sep 2026)
- 5. Anthropic Help Center: Does Anthropic crawl data from the web, and how can site owners block the crawler? (Accessed Sep 2026)
Cite this page
BrandWater AI Research. (2026, 21 September 2026). AI crawler types explained: training, search indexing and user-triggered fetching. https://brandwaterai.in/research/ai-crawler-types-explained
How we work
Figures are dated and linked to their sources. Where none exist we say so. Read our methodology and AI transparency pages.