Skip to content
BrandWater AI
IntermediateTechnical

AI crawlers and robots.txt: which bots to allow, and what each one does

Know the difference between training crawlers, search crawlers and user-triggered fetches before you edit robots.txt.

Time
4 min
Level
Intermediate
Updated
18 August 2026
Sections
8

Quick answer

AI companies run different crawlers for training, for building a search index and for fetching a page when a user asks. To appear in AI search answers, do not block the search crawlers. Whether to allow training crawlers is a business decision. Always confirm current names in each vendor's documentation.

Key takeaways

  • Separate three purposes: training, search indexing and user-triggered fetching.
  • Blocking a search crawler can remove you from that assistant's search answers.
  • Bot names change. Check the vendor documentation before editing robots.txt.

What kinds of AI crawlers exist?

Examples of documented crawlers (verify current names in the sources below)
VendorExample user agentDocumented purpose
OpenAIOAI-SearchBotSurfacing sites in ChatGPT search features
OpenAIGPTBotCollecting content that may be used to train models
OpenAIChatGPT-UserFetching a page when a user asks ChatGPT to
PerplexityPerplexityBotBuilding the Perplexity search index
PerplexityPerplexity-UserFetching a page when a user asks
GoogleGoogle-ExtendedA control token for whether content may be used for some Google AI models. Ordinary Googlebot handles Search.

How do you decide what to allow?

  1. 1List every bot in your server logs that identifies as an AI crawler.
  2. 2For each, read the vendor's page and record its stated purpose.
  3. 3Allow crawlers that build search indexes for platforms where you want to be visible.
  4. 4Decide separately about training crawlers, based on your content and legal position.
  5. 5Update robots.txt, then check your logs to confirm the change took effect.

An example robots.txt fragment

  • User-agent: OAI-SearchBot followed by Allow: / lets that search crawler read the site.
  • User-agent: PerplexityBot followed by Allow: / does the same for Perplexity.
  • User-agent: GPTBot followed by Disallow: / opts out of that training crawler only.

robots.txt is a request, not a lock

Reputable crawlers honour it. It does not protect private content. Use authentication for anything confidential.

The documented agents in one place

What each provider says its agents do
ProviderAgentDocumented purposerobots.txt
OpenAIOAI-SearchBotSurface websites in ChatGPT's search featuresDisallowing keeps a site out of ChatGPT search answers, though it can still appear as a navigational link; changes take about 24 hours
OpenAIGPTBotCrawl content that may be used to train foundation modelsDisallowing indicates content should not be used for training
OpenAIChatGPT-UserCertain user actions in ChatGPT and Custom GPTsUser-initiated, so robots.txt rules may not apply
PerplexityPerplexityBotSurface and link websites in Perplexity results; not used to train foundation modelsRespected; allowing it is recommended to appear in results
PerplexityPerplexity-UserVisit a page when a user asksGenerally ignores robots.txt because the fetch is user-requested
AnthropicClaudeBotCollect content that may contribute to trainingRestricting excludes future materials from training
AnthropicClaude-UserRetrieve pages for a person's questionDisabling may reduce visibility for user-directed search
AnthropicClaude-SearchBotImprove search result qualityDisabling may reduce visibility in search results

Google uses its own crawler for Search, and documents that robots.txt directives for Googlebot are the control for access. Names and behaviour change, so treat the provider pages listed in the sources as the authority.

Three decisions, not one

  1. 1Do you want to be eligible for search-connected answers? If yes, allow the search crawlers.
  2. 2Do you want your content used for model training? This is a business and legal decision, made separately.
  3. 3Do you need to keep certain content out of reach entirely? Use authentication, because robots.txt is a request and does not bind user-initiated fetches.

A worked robots.txt pattern

A site that wants to appear in search-connected answers but opt out of training could allow OAI-SearchBot and PerplexityBot and disallow GPTBot and ClaudeBot, then leave the rules for user-initiated agents to the provider documentation. The exact rules should be written and tested against each provider's current documentation before deployment. This describes a pattern and is not advice for any particular site.

Verify in your logs

  • Search your server or CDN logs for each user agent you have allowed or blocked.
  • Confirm blocked agents stop after the documented delay.
  • Confirm allowed agents can fetch your key pages and receive normal responses.
  • Check that a CDN or firewall rule is not blocking crawlers you meant to allow.

A policy your team can write down

Turn the three decisions into a short written policy: which search crawlers are allowed, what the position on training crawlers is and why, and how private content is protected. Give it an owner and a review date. Writing it down prevents the common situation where a robots.txt file accumulates rules nobody remembers adding.

Review the policy whenever a provider changes its documentation, and at least once a quarter. Crawler names, purposes and behaviour are not fixed, and a rule that made sense last year may block something you now want, or allow something you now would not.

Will allowing a crawler guarantee my brand is cited?

No. It only makes citation possible. Content quality, relevance and third party corroboration still decide whether you are named.

Can I block training but stay in AI search?

Vendors that publish separate crawlers for each purpose let you do this. Confirm the current behaviour in their documentation.

Where do I check this on my own site?

Open yourdomain/robots.txt and review your server logs for AI user agents.

How long after editing robots.txt does the change take effect?

It depends on the provider. OpenAI documents about 24 hours for OAI-SearchBot. Check each provider's documentation.

Can I block only some paths?

Yes. robots.txt supports path-specific rules. Keep them simple and test them.

Cite this page

BrandWater AI Research. (2026, 18 August 2026). AI crawlers and robots.txt: which bots to allow, and what each one does. https://brandwaterai.in/guides/ai-crawlers-and-robots-txt

How we work

Figures are dated and linked to their sources. Where none exist we say so. Read our methodology and AI transparency pages.

Put the theory to work on your brand.

Start with a first audit. We show where you appear, who is named instead, and what to fix first.