How AI engines pick sources: a working model of the pipeline
Every AI assistant with web search runs some version of the same four steps. Knowing where in that pipeline you are losing tells you what to fix.
By Shimon Carroll, Founder, SEO for AI Agents · Published
AI engines pick sources in four stages. First the assistant decides whether the question needs a web search at all. Then it rewrites the question into one or more searches. Then it retrieves candidate pages from a search index and narrows them down. Finally it writes an answer from the pieces it kept and attaches citations to the passages it used. A site can be lost at any stage: never searched for, never retrieved, retrieved but discarded, or read but not quoted. Each loss has a different fix, which is why "optimize for AI" is too vague to act on.
Stage 1: deciding whether to search
Assistants do not search for everything. Anthropic's documentation for Claude's web search tool is unusually explicit about the trigger: Claude searches when a request depends on information that is "current, changing, or outside its training data", such as current prices, recent announcements, or "information about specific organizations, people, or products that might have changed." It answers directly, without searching, for stable knowledge such as established facts and concepts.
This stage explains a pattern many brands notice: for general questions an assistant answers from memory and names whoever it learned about in training, while for questions about prices, availability, or a specific company it searches and cites. You influence the first case slowly, through how widely and consistently the web describes you over time; see why brand mentions matter. You influence the second case through everything below.
Stage 2: turning one question into several searches
The question a person types is rarely the search that runs. Google says AI Overviews and AI Mode use "query fan-out", issuing "multiple related searches across subtopics and data sources." Anthropic's documentation notes that a single request can trigger several searches, with simple factual questions using one to three and comparative research ten or more. Each of those searches is a separate chance to be found.
The practical consequence is that the page that ranks for the head term is no longer the only page in contention. In a March 2026 study of 4 million AI Overview citations, Ahrefs found only 37.9 percent of cited URLs ranked in the top 10 for the original query. The rest were found through the related searches. Content that answers the sub-questions, such as costs, comparisons, steps, and edge cases, gets more entries in the lottery.
Stage 3: retrieving and narrowing candidates
Each search runs against an index, and the engines do not share one. Google's AI features use Google's index. OpenAI runs its own search crawler, OAI-SearchBot, which it says is used "to surface websites in search results in ChatGPT's search features", and when Seer Interactive compared ChatGPT search citations with search results for the same 100 queries, more than 87 percent matched Bing's top results. Perplexity and Anthropic run their own search crawlers too. Being absent from an index, because a crawler was blocked or received an empty client-rendered page, ends your chances before any quality judgment is made.
From the retrieved candidates, the system keeps a handful. The vendors do not publish their selection rules, so this is the stage where honest analysis must stop at inference. What the evidence consistently points to is that traditional ranking quality carries through, since the candidates come from ranked search results, and that a few sources dominate each topic, a pattern our glossary calls the citation oligarchy.
Stage 4: composing the answer and citing passages
Finally the model writes the answer from the passages it kept and attaches citations. Anthropic's documentation shows what a citation actually contains: the URL, the page title, and a "cited_text" field holding up to 150 characters of the passage being cited, and it says citations are always enabled for web search. The unit being cited is a short piece of text, not a page.
This is the stage the GEO research measured. In experiments on 10,000 queries, adding quotations, statistics, and citations to source content raised its visibility in generated answers by up to about 40 percent, while keyword stuffing reduced it. Passages that state a clear, specific, attributable fact are easier to quote than passages that circle the point. Our guide to writing citable passages turns that into a template.
Two routes into an answer: memory and retrieval
The four stages describe retrieval, the route where an assistant searches and cites. The other route is memory: what the model learned during training and repeats without searching. The two are fed by different crawlers. Training crawlers such as GPTBot and ClaudeBot collect pages for future models; search crawlers such as OAI-SearchBot and Claude-SearchBot build the indexes that retrieval runs against. Blocking a training crawler affects the memory route for future models and leaves retrieval alone. Blocking a search crawler removes you from retrieval immediately. Our AI crawler directory lists which is which.
The memory route is slow to change and hard to measure from the outside, but it is not out of reach. Models learn about brands from the same public pages everyone else reads, so the work that helps retrieval, clear descriptions, consistent facts, and mentions on well-read sites, is also what future models will absorb. The difference is the time lag: a fix to retrieval can show up within days of a recrawl, while a change in what models remember waits for the next training run.
Where most sites lose
| Stage | Symptom | Usual cause | First fix |
|---|---|---|---|
| Decide | Assistants name competitors from memory, never you. | Thin presence in the sources models learn from. | Earn consistent third-party mentions and reviews. |
| Fan out | You rank for the head term but are not cited. | No content for the sub-questions. | Answer costs, comparisons, and steps on dedicated pages. |
| Retrieve | You are missing from one engine but present in others. | Crawler blocked, empty client-rendered HTML, or weak Bing visibility. | Check crawler access and server rendering; check Bing. |
| Compose | Your page is listed as a source but your words are not used. | No passage answers cleanly on its own. | Rewrite the first passage under each heading. |
One more failure sits outside the pipeline but hurts trust in it. In the Vercel and MERJ crawler study, 34.82 percent of ChatGPT's crawler fetches and 34.16 percent of Claude's hit 404 errors, partly from requesting URLs that do not exist. If an assistant links to a wrong or outdated URL on your site, a redirect from that path to the right page recovers the visit.
How SEO for AI Agents measures this
Our audit is organized around the same stages. For retrieval, the AI crawler access check confirms each engine's crawler is allowed in and actually receives a 200 response, and the AI crawler readability check confirms the content is in the HTML it receives. For composition, the passage extractability check maps which paragraphs on a page can be quoted cleanly.
For the outcome, we ask each engine the questions your buyers ask, record whether you were cited across repeated runs, and capture the sources it cited instead with the source of citation check. The engine divergence view puts the engines side by side, which is usually the fastest way to see which stage is failing: missing everywhere suggests the decide or fan-out stage, missing from one engine suggests retrieval. Every verdict links to the verbatim answer it came from.
Keep reading
Sources
- Anthropic, Claude web search tool documentation
Anthropic
- Google Search Central, AI features and your website (updated December 10, 2025)
Google Search Central
- OpenAI, Overview of OpenAI crawlers
OpenAI
- Seer Interactive, 87 percent of SearchGPT citations match Bing top results (February 6, 2025)
Seer Interactive
- Ahrefs, AI Overview citations and the top 10 (March 2, 2026)
Ahrefs
- Aggarwal et al., GEO: Generative Engine Optimization (KDD 2024)
arXiv / ACM KDD
- Vercel and MERJ, The rise of the AI crawler (December 17, 2024)
Vercel