Skip to content
BrandWater AI
Research note

AI crawler types explained: training, search indexing and user-triggered fetching

21 September 2026 3 min read

Headline findings

  • Three purposes: training, search indexing, and user-triggered fetching.
  • Vendors publish separate user agents for these purposes.
  • Confirm current names and behaviour before changing robots.txt.

Quick answer

Major AI vendors document separate crawlers for training data, search indexing and fetching a page at a user's request. Because the purposes differ, blocking one does not necessarily block the others. Names and behaviour change, so confirm them in the vendor documentation linked here.

Documented crawler purposes by vendor (as published, accessed Sep 2026)
VendorSearch indexTrainingUser-triggered
OpenAIOAI-SearchBotGPTBotChatGPT-User
PerplexityPerplexityBotNot stated as a separate crawlerPerplexity-User
GoogleGooglebot (Search)Google-Extended control tokenGoogle-specific fetchers

Method note

This table reports only what each vendor states in its own documentation on the access date. It makes no claim about crawl volume or frequency, and vendors can change names and behaviour at any time.

What does this mean for visibility?

A brand that blocks a vendor's search crawler may be left out of that vendor's search grounded answers. A brand that blocks a training crawler is making a separate decision. Treat the two independently.

Adding Anthropic and the user-initiated fetches

The table above lists OpenAI, Perplexity and Google. Anthropic documents three agents as well. ClaudeBot collects web content that may contribute to training. Claude-User supports Claude users: when a person asks a question, it may access websites. Claude-SearchBot navigates the web to improve search result quality. Anthropic explains that restricting ClaudeBot signals that future materials should be excluded from training datasets, while disabling Claude-User or Claude-SearchBot may reduce visibility in user-directed search and search results respectively.

The same three purposes across providers
PurposeOpenAIPerplexityAnthropic
Search indexingOAI-SearchBotPerplexityBotClaude-SearchBot
Model trainingGPTBotNot used to crawl for foundation models, per PerplexityClaudeBot
User-initiated fetchChatGPT-UserPerplexity-UserClaude-User

What the documentation says about robots.txt

  • OpenAI: disallowing OAI-SearchBot keeps a site out of ChatGPT search answers, though it can still appear as a navigational link, and changes take about 24 hours to process. Each setting is independent of the others.
  • OpenAI: because ChatGPT-User actions are initiated by a user, robots.txt rules may not apply.
  • Perplexity: PerplexityBot respects robots.txt and is recommended if you want to appear in results. Perplexity-User generally ignores robots.txt because a user requested the fetch.
  • Anthropic: each agent is a separate robots.txt token, and restricting one does not restrict the others.

Why this matters more than it looks

The most common mistake is treating all AI bots as one thing. A site that blocks everything to protect its content from training can also remove itself from search-connected answers. A site that allows everything may be making a decision about training use that it did not intend. Separating the three purposes lets each be decided on its own merits.

How to check your own setup

  1. 1Read your current robots.txt and list every AI-related user agent it mentions.
  2. 2Match each to the provider's documentation and write down its purpose.
  3. 3Decide the search, training and user-fetch settings separately for each provider.
  4. 4After changing rules, check server logs to confirm the behaviour changed after the documented delay.
  5. 5Review quarterly, because names and behaviour change.

robots.txt is a request

Reputable crawlers honour it, and user-initiated fetches may not follow it. It does not protect private content. Use authentication for anything that must stay private.
Where can I see which bots visit my site?

In your server or CDN logs. Filter by the user agent strings listed in each vendor's documentation.

Are these lists complete?

No. New crawlers appear, so treat the vendor pages as the source of truth.

Will allowing search crawlers expose my content to training?

The providers document search and training crawlers as separate agents with independent settings. Confirm each provider's current documentation, because policies can change.

What about crawlers I have never heard of?

Log them, check the provider's documentation, and treat unknown agents with caution. New crawlers appear regularly.

Cite this page

BrandWater AI Research. (2026, 21 September 2026). AI crawler types explained: training, search indexing and user-triggered fetching. https://brandwaterai.in/research/ai-crawler-types-explained

How we work

Figures are dated and linked to their sources. Where none exist we say so. Read our methodology and AI transparency pages.

Put the theory to work on your brand.

Start with a first audit. We show where you appear, who is named instead, and what to fix first.