AI Visibility pillar / check ai-citation-presence

AI citation rate and measurement stability (the Variance Engine)

AI engines are non-deterministic: ask the same question five times and the answer moves. We poll the identical prompt N times per engine (one to two runs today), count how often a brand is actually cited, publish a 95 percent Wilson confidence interval anyone can recompute, and flag results that are too volatile to optimize yet.

By Shimon Carroll, Founder, SEO for AI Agents · Last updated

What this check measures

For a given brand and query, we send the identical prompt to each configured AI engine N times (N is small and set by the plan: one run per engine on an on-demand check, two runs per engine on each weekly tracked-prompt refresh). For every poll we record whether the brand was actually cited in the answer, then compute two things over the successful polls. First, the observed citation rate: k cited out of n successful polls, reported as a plain k-of-n count and a rate. Second, a 95 percent confidence interval around that rate using the Wilson score interval, the standard interval that stays honest at small sample sizes and at rates near zero or one (our regime). We also assess how stable the result is across repeated runs and, when it is too noisy to act on, we say so explicitly. The denominator is always the number of polls that actually succeeded; a poll that timed out or hit a paused engine is never silently counted as "not cited." Each poll is an API sample: the engine's API model with its web-search tool, in the provider's default locale. Google AI Overviews is read from the Google results page through Serper, not an official Google API, and a run that returns no AI Overview block is reported as not measured. None of this reproduces the consumer apps, which add personalization, location and product orchestration we do not control.

Why it matters

A single confident-looking score is statistically misleading for a system that is non-deterministic by design. The same prompt sent to the same engine moments apart can cite a brand once and omit it the next time, so any tool that reports one number with no interval is hiding the uncertainty rather than measuring it. That is the gap this check closes. By polling repeatedly and publishing the confidence interval, we report what is actually known and how firmly it is known: "cited in 4 of 5 runs, 95 percent CI 38 to 98 percent" tells the truth that "80 percent visible" hides. The most valuable output is often the honest negative: when a result is too volatile to optimize yet, spending effort chasing it is premature, and we will tell you to wait for more signal rather than sell motion. This is the measurement-validity foundation the rest of the product stands on.

How we score it

We compute the observed citation rate as k cited out of n successful polls (zero when nothing succeeded), then the 95 percent Wilson score confidence interval around it. The Wilson interval is a published, standard statistical method: with z = 1.96, the interval is centered at (p + z squared / 2n) divided by (1 + z squared / n), with a half-width of z times the square root of (p(1 - p)/n + z squared / 4n squared) divided by (1 + z squared / n), clamped to the zero-to-one range. Anyone can recompute it by hand from the published k and n, and arrive at exactly the interval we show. Beyond the interval, we additionally assess how stable the result is across repeated runs and assign a coarse stability band, but we publish only the verdict (the band, and the volatility flag), never the internal scoring used to reach it. That separation is deliberate: the measurement is fully reproducible, the editorial scoring is not the point of the page.

Confidence-flag rules

Every measurement carries an explicit honesty flag. When the result is too volatile across repeated runs to be acted on, we raise the "too volatile to optimize yet" flag: the citation rate and interval are still shown, but we tell you plainly that the signal is too noisy to chase and that the right move is to wait for more measurement, not to optimize against noise. Small samples are treated with appropriate humility: a brand cited in 3 of 3 polls is not "rock solid" at 100 percent, because the Wilson interval at n = 3 is genuinely wide, and we report it that way rather than rounding up to false certainty. Some engines are near-deterministic rather than freshly resampled: Google AI Overviews and similar SERP-derived surfaces return what the search result page exposes, so polling them repeatedly measures our own parsing rather than the engine, and for those we report the citation interval but suppress the stability verdict and label the result single-source. When no poll succeeded, we report "not measured" rather than inventing a zero.

Common mistakes

  • Treating a single AI answer as ground truth. One poll of a non-deterministic engine is a sample of one, not a measurement; the brand could be cited or omitted on the very next identical prompt.
  • Reporting a point score with no interval. "80 percent visible" with no confidence range hides exactly the uncertainty that matters for a non-deterministic system.
  • Reading a small-sample 100 percent as certainty. Cited in 3 of 3 runs has a wide confidence interval; it is encouraging, not settled.
  • Optimizing against a too-volatile result. When the signal is flagged too volatile to optimize yet, chasing it burns effort on noise; the honest move is to wait for more measurement.
  • Counting failed or timed-out polls as "not cited." The honest denominator is the number of polls that actually succeeded, never the number attempted.

How to fix it

There is nothing to "fix" in this check the way a missing tag is fixed; it is a measurement, and the response depends on what it reports. When the citation rate is low with a tight interval, that is a real, stable visibility gap worth working: the other AI-visibility checks on this report (crawler readability, schema completeness, passage extractability, entity graph) name the concrete levers. When the result is flagged too volatile to optimize yet, the correct action is to wait and re-measure rather than react, because the signal is not yet stable enough to attribute any change to your work. When you do act, re-run the measurement afterward and compare intervals, not point estimates, so you can tell a genuine improvement from run-to-run noise. Every measurement links to the verbatim AI answer, the exact model, and the timestamp behind it, so any number on the report can be opened and verified one click away, and the Wilson interval recomputed from the published k and n by anyone who wants to check our arithmetic.

Primary sources

Changelog

  • · Sample counts corrected to the shipped plans: one run per engine on an on-demand check, two per weekly refresh. Every engine result is an API sample (Google AI Overviews: SERP-derived), not a reproduction of the consumer app.
  • · Initial publication. Documents the Variance Engine: repeated-poll sampling, the published Wilson 95 percent confidence interval, one-click receipt verification, and the "too volatile to optimize yet" honesty flag.