Technical pillar / check root-response-access

Page loads as content: did our crawler get your real page?

We check whether the page you asked us to audit answered our crawler with a readable web page, or with an error, a refusal, a rate limit, something that is not a web page, or a bot check. When it did not, we grade nothing about your content and tell you what the server sent instead.

By Shimon Carroll, Founder, SEO for AI Agents · Last updated

What this check measures

The HTTP status our crawler received for the audited page, its content type, and whether the body is a bot check or challenge page rather than your site. A 401 or 403 (access refused), a 429 (rate limited), any other error status, a body that is not HTML, or a page whose title or markup is a known challenge (for example a "Just a moment..." screen) counts as not content. When we see one, the finding quotes the status, the content type and the challenge text we matched.

Why it matters

Google's Search Essentials list three technical requirements for a page to be eligible for Search: Googlebot is not blocked, the page works (Google receives an HTTP 200 success status), and the page has indexable content. Google's SEO Starter Guide also tells site owners to check whether Google can see the page the same way a user does. Bing's Webmaster Guidelines ask sites to allow Bingbot to crawl and render content efficiently, list blocking important URLs unnecessarily among the things to avoid, and say content that cannot be reliably rendered may not be indexed or selected for grounding results. A bot wall that stops our crawler often stops AI crawlers too, and it means every content finding we could report would describe the wall, not your business.

How we score it

Measured, once per audit. A readable HTML page passes. A refusal, rate limit, error page or challenge is a MEDIUM finding; a body that is not HTML is LOW. On a page that did not load as content we do not run the content checks at all: trust pages, people, reviews, headings, writing and data checks are each shown as not measured, with the response receipt, and none of them changes your score. Checks that read the response or your site data (robots.txt, sitemap, HTTPS, server speed, crawler access, Search Console data) still run. A challenge script or a CAPTCHA on an ordinary full page does not count as a challenge.

Confidence-flag rules

MEDIUM. The status code and content type are exact. Whether a search engine gets the same answer depends on your firewall rules: many bot walls let verified Googlebot and Bingbot through while blocking other crawlers, so this finding proves what our crawler received, not what every engine receives.

Common mistakes

  • Turning on a strict bot-fight or challenge mode at the CDN and forgetting it also challenges search and AI crawlers.
  • Rate limiting by IP so tightly that a crawler fetching a handful of pages is cut off.
  • Serving a login or age gate on the homepage with no indexable content behind it.

How to fix it

Open your CDN or firewall bot settings and allow the crawlers you want to be found by (verified search engine bots and the AI crawlers you accept), or lower the challenge level for your public pages. Make sure the homepage returns HTTP 200 with your real content, then re-run the audit.

Primary sources

Changelog

  • · Initial publication. Pages that answer with a bot wall, a challenge, an error or a non-HTML body are no longer graded as your content; every content check on them is marked not measured.