Schema markup for AI citation: the types that matter
Structured data is no longer mostly about rich results. In the AI era it is how engines disambiguate your entity and decide whether to trust and cite you.
By Shimon Carroll, Founder, SEO for AI Agents · Published
For years, the case for schema markup was rich results: review stars, FAQ accordions, product prices in the search listing. That case still holds, but it is no longer the most important one. In the AI era, the primary job of structured data is entity disambiguation, telling search and AI engines exactly what your page and your brand are, so they can connect you to the right node in their knowledge graph and decide whether to trust and cite you.
This reframing changes which types matter and why. Below is the practical set, what each does, and the entity work that makes all of them more powerful. Notably, the legacy tool we benchmark against ships zero JSON-LD across its own pages, which is a striking gap for a company that sells optimization. Doing schema right is a real, available advantage.
Format first: JSON-LD, server-rendered
Use JSON-LD, the format Google recommends, in a script block of type application/ld+json. It keeps the markup separate from your presentation, is easy to generate from data, and is straightforward to validate. The one non-negotiable: render it server-side. If your schema only appears after JavaScript runs, the AI crawlers that do not execute JavaScript will never see it, the same readability problem that hides your content hides your markup.
The types that matter, and what each does
Organization and WebSite (homepage)
These declare who you are. Organization names the entity, its logo, and, critically, its sameAs links, the URLs of your authoritative profiles (LinkedIn, Crunchbase, your Wikidata entry, official social accounts). The sameAs array is what ties your on-site entity to the off-site graph, which is how an engine confirms you are a recognized thing it can cite confidently. WebSite enables the sitelinks search box and names the site as an entity.
Article (editorial content)
Article markup carries authorship and dates: the headline, the author as a Person (ideally with their own sameAs links), the published and modified dates, and the publisher. For consequential topics, named, credentialed authorship is an E-E-A-T and AI-trust signal. Dates establish freshness. This is why our own glossary, methodology, and blog content ships Article and author markup, the byline is part of the trust argument.
Product and Offer (commerce)
Product describes what you sell, name, description, brand, and Offer carries price, currency, and availability. AggregateRating, when you can genuinely substantiate it, adds review signal. For AI shopping and recommendation queries, this markup is how an engine knows your product is a specific, well-defined offering rather than an ambiguous string, which is the difference between being recommended and being skipped.
LocalBusiness and its subtypes (physical businesses)
For a physical or service-area business, LocalBusiness restates your NAP, hours, and geo-coordinates in machine-readable form, reinforcing the entity across your citations and Google Business Profile. Use the most specific accurate subtype, LegalService for a law firm, MedicalBusiness for a practice, AccountingService for a CPA, RealEstateAgent for an agent, because specificity is what helps an engine place you correctly. Generic LocalBusiness loses to a specific subtype.
FAQPage and HowTo (extractable answers)
FAQPage marks up question-and-answer pairs; HowTo marks up step-by-step instructions. Google has narrowed the rich-result display for FAQ to a small set of authoritative sites, so the SERP enhancement is no longer guaranteed, but the structured form still helps engines parse your answers, and it pairs naturally with the passage structure that wins AI citation. Mark up answers you actually show, and structure the visible content the same way: a clear question, a tight answer.
BreadcrumbList (structure)
BreadcrumbList describes a page's position in the site hierarchy. It is cheap, it powers breadcrumb display, and it gives engines an explicit map of how your content is organized, which reinforces topical relationships. We ship it on every editorial and comparison page.
The work that makes schema actually pay off
Markup alone is necessary but not sufficient. Three disciplines turn schema from decoration into citation leverage:
- Mark up what is actually on the page. Google penalizes structured data that describes content the user cannot see, or that misrepresents the page. The markup must match the visible content exactly.
- Use the most specific accurate type, and validate against the live spec. Schema.org is a living vocabulary and Google's rich-result requirements evolve; validating against a 2020 snapshot will let stale markup through. Validate against current.
- Connect to the entity graph with sameAs. The single highest-leverage addition is linking your Organization and Person entities to their authoritative external profiles, above all a Wikidata entry. Wikidata has a lower notability bar than Wikipedia and gives the engines a machine-readable node to anchor your identity to.
Schema makes a page eligible for rich results and, more durably, tells search and AI engines what your entities are. The second job is the one that compounds.
SEO for AI Agents, methodology
A pragmatic rollout
You do not need every type at once. A sensible order for most sites: Organization and WebSite on the homepage with full sameAs first, because that establishes the brand entity that everything else leans on. Then Article on your editorial content, Product and Offer on commerce pages, LocalBusiness with the right subtype if you are a physical business, and BreadcrumbList everywhere. Add FAQPage and HowTo where the content genuinely fits.
Then do the entity work upstream of the markup: a Wikidata entry, consistent facts across the sources the engines trust, named authors with real credentials. Schema points at your entity; the entity has to exist for the pointing to mean anything.
Our audit validates your structured data against the live Schema.org spec, checks for the entity-clarifying types and the sameAs connections, and flags markup that does not match the page. It also audits the negative space, the high-value types you are missing, because for AI citation, the schema you have not shipped is often the gap between being recognized and being ignored.
Keep reading
Sources
- Google Search Central, intro to structured data
Google Search Central
- Schema.org, full type hierarchy
Schema.org
- Google, how the Knowledge Graph works
Google
- Wikidata, introduction
Wikimedia Foundation