AI Visibility pillar / check ai-crawler-readability

AI crawler readability (raw HTML token ratio)

A page should expose at least 70 percent of its rendered DOM token count in the raw HTML, since GPTBot, ClaudeBot, and PerplexityBot do not execute JavaScript.

By Shimon Carroll, Founder, SEO for AI Agents · Last updated

What this check measures

For each page, we compute the token count of the raw HTML body (after stripping script and style content) and the token count of the JS-rendered DOM body. The ratio (raw / rendered) is the AI crawler readability score. We additionally verify that every H1 and H2 in the rendered DOM appears in the raw HTML; the heading subset is a sharper signal than aggregate token count.

Why it matters

AI crawlers (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, ChatGPT-User, Google-Extended, Bytespider, CCBot) are documented to fetch raw HTML without executing JavaScript in their production paths (per OpenAI, Anthropic, and Perplexity's own published crawler documentation as of late 2025). A page where the primary content is hydrated client-side is largely invisible to these crawlers. Across our launch-vertical audits, pages with a readability score below 70 percent have a near-zero cited-answer rate on ChatGPT, Claude, and Perplexity; pages above 90 percent are cited at approximately the rate predicted by their schema density and entity-graph completeness.

How we score it

PASS if (raw HTML body tokens / rendered DOM body tokens) >= 0.70 AND every H1 and H2 in the rendered DOM is present in the raw HTML. FAIL on either condition. Severity defaults to HIGH for the AI Visibility pillar. The ai_agencies adapter does not bump severity (already HIGH); MCA, lawyers, doctors, and real_estate adapters apply the default.

Confidence-flag rules

Confidence is HIGH when both the raw fetch and the Playwright render completed successfully and returned 2xx. Confidence drops to MEDIUM when the Playwright render exceeded 10 seconds (the comparison may not reflect what a normal-budget render would see). Confidence is LOW when either fetch returned 4xx or 5xx; the finding is suppressed in that case.

Common mistakes

  • Assuming that "Google can render JavaScript" generalizes to AI crawlers; it does not.
  • Building a Next.js App Router page entirely with client components, leaving the raw HTML body containing only the chrome and an empty root.
  • Loading primary content with a useEffect fetch after page load; the raw HTML contains no answer to the buyer's question.
  • Using a third-party widget for the main hero or feature grid that renders client-side, with no static fallback in the initial HTML.

How to fix it

Move primary content out of client components and into Server Components, static export, or generateStaticParams. Verify with a curl that does not execute JavaScript ("curl -s https://yoursite/" piped to grep for expected text). For Next.js App Router, the default Server Component path renders to HTML; convert any unnecessary "use client" boundaries back to Server Components. For SPAs (Vite, CRA, Remix client-only), introduce a server-render layer (Vite SSR, Remix Server) or a static fallback shell.

Primary sources

Changelog

  • · Initial publication. Threshold set at 70 percent based on observed correlation with cited-answer rate across 200 launch-vertical audits.