AI Visibility pillar / check ai-crawler-access
AI crawler access: can AI engines actually fetch you?
An AI engine that cannot fetch your page cannot cite it. For each major AI bot we check both what robots.txt allows and what your site actually serves when that bot knocks, because a CDN or firewall can block a crawler your robots.txt never mentioned.
By Shimon Carroll, Founder, SEO for AI Agents · Last updated
What this check measures
We evaluate access for the major AI crawlers by their published User-Agent strings, including OpenAI (GPTBot, OAI-SearchBot, ChatGPT-User), Anthropic (ClaudeBot, Claude-User), Perplexity (PerplexityBot, Perplexity-User), Google-Extended, Common Crawl (CCBot), Amazonbot, Bytespider, and Meta-ExternalAgent. The check has two layers. First, for every bot we read your robots.txt and report whether it is allowed to fetch the page and which exact directive decided that, a verbatim allow or disallow line. Second, for a small high-value subset we go further and actually fetch the page, once each, sending the bot real User-Agent, and observe what your server returns. That bounded live probe is what catches a block robots.txt never declared: a firewall or CDN returning a 403 forbidden, a 429 throttle, or a JavaScript challenge to a specific bot User-Agent. We capture the real HTTP status and selected response headers as proof.
Why it matters
AI visibility starts with a fetch. If GPTBot or ClaudeBot cannot retrieve your page, that engine simply has no copy of your content to draw on, so it cannot cite you no matter how good the page is. The trap is that robots.txt is only half the story. A site can allow every AI bot in robots.txt and still hard-block them at the edge, because a security product or CDN is sniffing User-Agents and serving bot traffic a 403 or a challenge page that a normal browser never sees. That kind of block is invisible to any audit that only reads robots.txt, and it silently removes you from the engines entirely. By fetching as the bot actually fetches, we surface the gap between what you think you allow and what your infrastructure really serves, which is exactly the gap that quietly costs brands their AI presence.
How we score it
The robots layer is pure, public fact: anyone with your robots.txt can reproduce the same per-bot allow or deny verdict and see the same directive that decided it. The live layer reports the verbatim HTTP status and headers each bot User-Agent received, which anyone can reproduce with a single fetch sending the same User-Agent. From those facts we classify each probed response into a plain category, served, forbidden, rate limited, or challenge, derived directly from the status and headers, with no proprietary scoring. The only judgment in this check is how the finding severity reflects how many high-value bots are blocked, and that judgment lives in the engine, never in the evidence. The receipts, the directives and the raw responses, are everything you need to verify the verdict yourself.
Confidence-flag rules
Confidence is highest when the live probes ran and we observed real access, which happens on the audit root page. The probes are deliberately bounded for cost and politeness: they run on the root page only, never on deep pages, and only for a small high-value set of bots, fetched one at a time with a short timeout and no retries, obeying the same one-request-per-second courtesy the rest of the audit uses so the probe itself never provokes the block it is trying to measure. On deep pages, or when a probe times out, we fall back to the robots-only verdict at lower confidence and say so, rather than inventing a live result. A probe that fails to connect is reported as undetermined, not as a block, so we never accuse your site of a refusal we did not actually observe.
Common mistakes
- Assuming robots.txt is the whole story. A CDN or firewall can block GPTBot or ClaudeBot with a 403 or a challenge while robots.txt cheerfully allows them.
- Turning on aggressive bot protection that lumps legitimate AI crawlers in with scrapers, quietly removing the site from the engines that drive AI citations.
- Disallowing a bot in robots.txt by accident, often a copied template that blocks a User-Agent the brand actually wants to be cited by.
- Testing access from a browser and concluding all is well, when the server treats the bot User-Agent completely differently from a browser.
How to fix it
Make sure the AI bots you want to be cited by are explicitly allowed in robots.txt, then confirm your CDN or firewall is not blocking those same User-Agents at the edge. If a probe shows a 403, 429, or challenge for a bot you intend to allow, add an allow-rule for that User-Agent in your security or CDN configuration. Be deliberate, since allowing a training bot like GPTBot or Google-Extended is a content-usage decision as well as a visibility one, so allow the bots whose engines you want to appear in and document the choice. After you adjust robots.txt and the edge, re-run the audit and confirm the previously blocked bots now return a normal served status. Every verdict here is backed by a real directive and a real HTTP response, so you can confirm each one with your own fetch.
Primary sources
Changelog
- · Initial publication. Documents the two-layer access read: a per-bot robots.txt allow/deny matrix for every major AI crawler, plus a bounded live fetch on the root page using each high-value bot User-Agent to detect WAF/CDN blocks that robots.txt cannot reveal.