Technical SEO

Indexing

The process by which a search engine analyzes a crawled page and stores it in its index so it can be returned for queries.

By Shimon Carroll, Founder, SEO for AI Agents · Last updated

Indexing is the step between crawling and ranking. After a crawler fetches a page, the engine processes it, understands its content and canonical, and decides whether to store it in the index. Only indexed pages can rank. A page can be crawled but not indexed if the engine judges it low-quality, duplicate, thin, or explicitly excluded.

The controls are precise. A noindex meta tag or X-Robots-Tag header keeps a page out of the index even though it can still be crawled. A robots.txt disallow prevents crawling but does not reliably prevent indexing, which is the most common confusion: to remove a page from the index you must let it be crawled so the engine can see the noindex. Canonical tags consolidate duplicates onto one indexed URL.

The Index Coverage report in Search Console is the diagnostic surface: it shows which pages are indexed, which are excluded, and why (crawled-not-indexed, discovered-not-indexed, duplicate, soft 404, blocked). Reconciling that report against your sitemap and your intent is core technical hygiene. We diff intended-indexable pages against what Search Console reports as actually indexed.

Primary sources