Technical SEO
robots.txt
A file at the site root that tells crawlers which paths they may or may not fetch. In 2026 it is also where you make the allow-or-block decision for AI crawlers.
By Shimon Carroll, Founder, SEO for AI Agents · Last updated
robots.txt is a plain-text file at the root of a domain that gives crawlers per-user-agent rules about which paths to fetch. It is governed by the Robots Exclusion Protocol, which became an IETF standard (RFC 9309) in 2022. It controls crawling, not indexing: a disallowed page can still be indexed if other pages link to it, so to keep a page out of the index you use a noindex meta tag or header, not a robots disallow.
In 2026 robots.txt carries a new strategic weight, because it is where you decide how AI crawlers are treated. GPTBot, ClaudeBot, PerplexityBot, Google-Extended, OAI-SearchBot, ChatGPT-User, Bytespider, and CCBot can each be allowed or disallowed by name. This is a genuine choice with tradeoffs: blocking the AI crawlers protects your content from training use but also removes you from the data that lets those systems recognize and recommend your brand.
The common errors are blocking resources the page needs to render (CSS, JS) which breaks how Google sees the page, accidentally disallowing the whole site with a stray rule, and conflating crawl-blocking with index-blocking. We audit who is allowed, who is blocked, whether the choice is deliberate, and whether any rule is inadvertently hiding content from the engines you want to reach.
Related terms
Primary sources
- Google Search Central, robots.txt introduction
Google Search Central
- IETF, RFC 9309 Robots Exclusion Protocol
IETF