Technical pillar / check bing-crawl-access

Bing crawl access: can Bing read your site?

We read your robots.txt the way Bing does (a bingbot group, else msnbot, else the catch-all group) and check that bingbot may fetch your homepage and the audited page, that no bingbot-only noindex is set, and that any Crawl-delay is within what Bing recommends.

By Shimon Carroll, Founder, SEO for AI Agents · Last updated

What this check measures

From the robots.txt we already fetch for the audit, we pick the group Bing obeys: a group naming bingbot, otherwise one naming msnbot (Bing still reads it), otherwise the catch-all User-agent: * group. We evaluate your homepage and the audited page with the same rule matcher our crawler uses, read any Crawl-delay in that group, and look for noindex or none scoped to bingbot or msnbot in a meta tag or the X-Robots-Tag header. The evidence shows the matched group lines verbatim.

Why it matters

Bing powers Bing search and the answers Microsoft Copilot grounds on the web. Bing's Webmaster Guidelines list blocking Bingbot in your robots.txt file among the things to avoid, and say content that cannot be reliably crawled may not be indexed or selected for grounding results. A third-party study by Seer Interactive found that 87 percent of SearchGPT citations matched Bing's top results, which is one reason Bing access matters beyond Bing itself. Bing reads Crawl-delay as a throttle; its own guidance recommends against any value higher than 10. Google ignores Crawl-delay.

How we score it

Measured. A Disallow that keeps bingbot off your homepage or the audited page, or a bingbot-scoped noindex, is a HIGH finding: Bing cannot use the page. A Crawl-delay above the value Bing recommends is LOW. Otherwise the check passes. If you block Bing on purpose, ignore the finding. It runs once per audit because robots.txt is one file per host. If robots.txt does not load during the audit (a server error, rate limit or timeout), the robots half is not measured and the finding says so; a bingbot noindex on the page is still reported. A robots.txt that returns 404 means no rules, so Bing may crawl.

Confidence-flag rules

HIGH. robots.txt rules are deterministic and we show the exact group lines that decided the verdict. We cannot see firewall or CDN rules that block Bing by IP or user agent; the AI crawler access check covers live blocking for AI bots.

Common mistakes

  • Copying a staging robots.txt with Disallow: / to production.
  • Adding a bingbot group to slow Bing down and forgetting it overrides every rule in the catch-all group.
  • Setting a large Crawl-delay to save server load, which leaves Bing and its AI answers with stale copies of your pages.

How to fix it

Remove the Disallow that applies to bingbot, or narrow it to the paths you really want hidden. Remove a bingbot noindex from pages you want in Bing and its AI answers. Lower Crawl-delay to 10 or less, or drop it; Bing adjusts its own crawl rate.

Primary sources

Changelog

  • · An unreadable robots.txt (server error, rate limit or timeout) is reported as not measured instead of allowed. A bingbot noindex on the page is still reported.
  • · Initial publication. Measured check, run once per audit. Sources re-read in a real browser on 2026-10-07.