AI crawler directory 2026: every bot, what it does, what blocking costs
Every major AI company now runs separate bots for training, for search, and for live user requests. Blocking the wrong one removes you from AI answers; blocking the right one only opts you out of training.
By Shimon Carroll, Founder, SEO for AI Agents · Published
Most AI companies now run three kinds of bot, and they do different things. A training crawler, such as GPTBot or ClaudeBot, collects pages for future models. A search crawler, such as OAI-SearchBot, Claude-SearchBot, or PerplexityBot, builds the index those assistants cite from. A user fetcher, such as ChatGPT-User or Claude-User, visits a page live because someone asked about it. Block a training crawler and you opt out of training. Block a search crawler or user fetcher and you disappear from that assistant's answers. That distinction is the whole game, and the directory below is organized around it.
Every entry was checked against the operator's own documentation on October 4, 2026. Where an operator says its bot may not follow robots.txt, we say so.
The directory
| User agent | Operator and role | Obeys robots.txt | If you block it |
|---|---|---|---|
| GPTBot | OpenAI. Training: used "to make our generative AI foundation models more useful and safe." | Yes | Future OpenAI models do not train on your pages. ChatGPT search is unaffected. |
| OAI-SearchBot | OpenAI. Search: surfaces sites "in search results in ChatGPT's search features." | Yes. OpenAI says changes take about 24 hours to apply. | You stop appearing as a cited source in ChatGPT search. |
| ChatGPT-User | OpenAI. User fetcher for actions a ChatGPT user initiates. | Not necessarily. OpenAI says robots.txt "may not apply" to user-initiated requests. | Unreliable as a block. Use server rules if you must. |
| ClaudeBot | Anthropic. Training: collects content to improve its models. | Yes, including Crawl-delay | Your future content is excluded from Anthropic training data. |
| Claude-SearchBot | Anthropic. Search: indexes content to improve search results. | Yes | Anthropic says this "may reduce your site's visibility" in Claude's answers. |
| Claude-User | Anthropic. User fetcher when someone asks Claude a question. | Yes | Claude cannot retrieve your page in response to a user query. |
| PerplexityBot | Perplexity. Search: surfaces and links sites; Perplexity says it is not used to train foundation models. | Yes | You drop out of Perplexity's index and its cited sources. |
| Perplexity-User | Perplexity. User fetcher for live answers. | No. Perplexity says it "generally ignores robots.txt rules." | Not blockable via robots.txt. |
| Googlebot | Google. Search indexing, which also feeds AI Overviews and AI Mode. | Yes | You leave Google Search, including AI Overviews. Never block it. |
| Google-Extended | Google. A control token, not a separate crawler: governs use in Gemini training and grounding. | Yes | Opts out of Gemini model use. Google says it is not a Search ranking signal. |
| Applebot-Extended | Apple. A control token for Apple AI training; Apple says it "does not crawl webpages." | Yes | Opts out of Apple AI training. Apple search via Applebot is unaffected. |
| Meta-ExternalAgent | Meta. Crawls for AI training and product indexing. | Yes | Opts out of Meta AI training. |
| Meta-ExternalFetcher | Meta. User-initiated fetches. | Not necessarily. Meta says it "may bypass robots.txt rules." | Unreliable as a block. |
| Amazonbot | Amazon. Improves products and "may be used to train Amazon AI models." | Yes | Opts out of Amazon training. Amzn-SearchBot and Amzn-User are separate and do not train. |
| CCBot | Common Crawl. Builds an open web archive that many open-source models train on. | Yes | You leave future Common Crawl snapshots, and the models built from them. |
| DuckAssistBot | DuckDuckGo. Live fetches for its cited AI answers; DuckDuckGo says the data is not used to train models. | Yes, effective after 72 hours | You stop appearing in DuckDuckGo's AI-assisted answers. |
Two corrections to advice that still circulates. First, GoogleOther is not the AI Overviews crawler. Google describes it as a generic crawler used by product teams for things like research crawls; AI Overviews are built from the normal Search index that Googlebot crawls. Second, blocking Google-Extended does not remove you from AI Overviews. Google states that Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal." If you want out of AI Overviews specifically, the documented controls are snippet controls such as nosnippet, not Google-Extended.
Which bots should you allow?
For most businesses that want customers to find them through AI answers, the answer is all of the search crawlers and user fetchers, always. Training crawlers are a real choice. Allowing them may help future models know your brand; blocking them protects content you consider proprietary. Neither choice affects whether today's assistants can cite you, because citation runs through the search and user agents. Remember too that blocking training does not remove anything already collected; it applies to future crawls.
Template 1: maximum AI visibility
User-agent: *
Allow: /
Disallow: /admin/
Sitemap: https://example.com/sitemap.xmlYou do not need to list each AI bot by name if your wildcard group allows them. Listing bots only matters when you want them treated differently.
Template 2: opt out of training, stay in AI search
# Training crawlers and control tokens: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Amazonbot
User-agent: CCBot
Disallow: /
# Everyone else, including OAI-SearchBot, Claude-SearchBot,
# Claude-User, PerplexityBot, DuckAssistBot, and Googlebot
User-agent: *
Allow: /
Disallow: /admin/
Sitemap: https://example.com/sitemap.xmlThe robots.txt rule that trips up careful teams
Under the Robots Exclusion Protocol, standardized as RFC 9309 in 2022, a crawler that finds a group naming it obeys that group and ignores the wildcard group entirely. The wildcard only applies to crawlers no group names. So if you write a group for GPTBot that disallows one folder, GPTBot ignores every rule in your wildcard group, including disallows you assumed applied to everyone. When you add a named group for any bot, repeat every rule you want that bot to follow.
robots.txt is not the only gate
Three things can block an AI crawler that your robots.txt never mentions.
- CDN defaults. On July 1, 2025 Cloudflare announced it was changing its default to block AI crawlers for new domains unless the owner chooses otherwise. Many sites now block AI bots at the edge without anyone on the marketing team knowing.
- Firewall and bot-management rules that return a 403 error, a 429 rate limit, or a JavaScript challenge to unfamiliar user agents. A challenge page is a block for a crawler that does not run JavaScript.
- Rendering. A bot that is allowed in but receives a client-rendered shell has been let into an empty room. See which rendering modes AI crawlers can read.
The opposite problem exists as well. In August 2025 Cloudflare reported that Perplexity, when its declared crawler was blocked, fetched pages with an undeclared crawler that presented as a generic Chrome browser, at 3 to 6 million requests a day against 20 to 25 million for the declared crawler. Perplexity disputed the characterization. The practical takeaway is narrow: robots.txt expresses a preference, and if you need a hard block, enforce it at the server or CDN.
How SEO for AI Agents measures this
The AI crawler access check evaluates each named AI agent against your robots.txt and reports the exact directive that decided it, quoted from your file. For a small set of high-value agents it then sends one real request with that agent's published User-Agent and records the HTTP status your server returns, which is how we catch a CDN 403 or a challenge page that robots.txt never declared. The receipt is the status code and the directive, both of which you can reproduce with curl.
Access is necessary but not sufficient, so the same audit runs the AI crawler readability check on what those bots receive. Read the glossary entries for GPTBot, ClaudeBot, and PerplexityBot for each operator's history and documentation. We update this directory when an operator changes its documentation; the date at the top of the page tells you when it was last verified.
Keep reading
Sources
- OpenAI, Overview of OpenAI crawlers
OpenAI
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler
Anthropic
- Perplexity, Perplexity crawlers
Perplexity
- Google Search Central, Google's common crawlers (Google-Extended, GoogleOther)
Google Search Central
- Apple, About Applebot
Apple
- Meta, Meta web crawlers
Meta
- Amazon, Amazonbot and Amazon crawlers
Amazon
- Common Crawl, CCBot
Common Crawl
- DuckDuckGo, DuckAssistBot
DuckDuckGo
- Cloudflare, Perplexity is using stealth, undeclared crawlers (August 4, 2025)
Cloudflare
- Cloudflare, Content Independence Day: no AI crawl without compensation (July 1, 2025)
Cloudflare
- IETF, RFC 9309: Robots Exclusion Protocol (September 2022)
IETF