AI Visibility pillar / check source-of-citation

Source of citation: which sources actually drive your AI mentions

When an AI engine cites you, it cites you through a source. We collect the real URLs each engine returned and group them by domain, so you can see exactly which sources are feeding your AI mentions, instead of guessing.

By Shimon Carroll, Founder, SEO for AI Agents · Last updated

What this check measures

When an AI engine answers a question, it generally points at sources: the pages and domains it drew on to compose the answer. For each engine measurement, we capture the source references the engine actually returned, then group those references by their domain so you can see which domains appear most often behind your mentions. The output is a list of domains, and for each domain we show which engines surfaced it, how many of the captured citations it accounts for, and concrete example URLs you can open. The source URLs are the receipts: they are the real references the engine returned, not our inference about where an answer might have come from. When an engine returns an answer with no source references at all, we say so honestly rather than fabricating a source list.

Why it matters

Knowing that an engine cites you is the start; knowing what it cites is what you can act on. If most of your Perplexity mentions trace back to a single review directory, that directory is a lever you can pull, and a dependency you should understand. If your own domain rarely appears among the sources behind your mentions, that is a direct signal that the engines are citing other people talking about you rather than you talking about yourself. The source graph turns "you are visible" into "you are visible because of these specific sources," which is the difference between a vanity reading and a work plan. Because the underlying URLs are real and openable, every conclusion here is checkable rather than asserted.

How we score it

We group the captured source references by domain and present the grouped result. The grouping involves tidying up URLs so that the same page is not double counted because of trailing slashes, tracking parameters, or subdomain variants, and then organizing the domains so the most relevant sources surface first. The internal logic that decides how to organize and rank the domains is editorial and is deliberately not published, because it is not a measurement anyone needs to reproduce. What is reproducible is the receipt layer: the actual source URLs and the per-engine attribution counts are exposed, so you can verify which domains the engines returned and how often, and reach the same grouping by domain yourself.

Confidence-flag rules

The graph is only as complete as the sources the engines return. Some engines expose rich source references; others return few or none, and a SERP-derived surface reflects what the search result page already lists rather than an independent set of references. We never paper over that: when an engine returns no usable sources for a measurement, that engine contributes nothing to the graph and is shown as not disclosed for that run, rather than being filled in with a guess. URLs are tidied up defensively so that the same underlying page is counted once, but we never merge distinct pages or distinct domains to inflate a domain's apparent share. Example URLs shown for a domain are always real references the engine actually returned, capped to a representative few, and openable so you can confirm them yourself.

Common mistakes

  • Assuming an AI mention comes from your own site. Often the engine is citing a third party talking about you, which the source graph makes visible.
  • Treating "no sources returned" as "no sources exist." Some engines simply do not expose references; we mark that as not disclosed rather than inventing one.
  • Double counting the same page across URL variants. The same page behind trailing slashes, tracking parameters, or subdomains is one source, not several.
  • Optimizing a domain you cannot influence while ignoring one you can. The graph shows which sources are actually feeding your mentions so effort lands where it moves the needle.

How to fix it

Use the graph to find the sources doing the work and the gaps worth closing. When a third-party domain drives a large share of your mentions, the move is to strengthen and keep accurate the presence you have there, and to understand the dependency it represents. When your own domain is largely absent from the sources behind your mentions, that points back at the on-page AI-visibility checks (crawler readability, schema completeness, passage extractability) so the engines can cite you directly rather than only through others. After you act, re-run the measurement and compare which domains appear, because a healthier graph shows up as your own and your strongest sources gaining presence over time. Every domain in the graph is backed by real, openable URLs, so any claim about where your mentions come from can be verified one click away.

Primary sources

Changelog

  • · Initial publication. Documents the source-of-citation graph: collecting the actual source URLs each engine returned, grouping them by domain, and exposing those URLs as reproducible receipts while keeping the grouping and source-ranking logic server-side.