Finding · 2026-08-24

Their robots.txt says yes.
Their edge says no.

5,144 of 28,874 measured domains publish a robots.txt permitting a named AI crawler, and then refuse that same crawler at the network edge. Nobody decided this. It is what a bot-management default does to a policy nobody re-read.

Named, and checked minutes before publication 30 of 30 re-probed live

Each row is two requests to the site root, made from outside the target's network: a control carrying an ordinary browser user agent, and one carrying the crawler's published user agent. The control returned a page. The crawler did not. The date on each row is when that pair of requests was last made.

This list is generated from the corpus on each page load rather than frozen when the page was written. About 1% of measured domains change crawler policy per week, so a fixed list would keep naming companies that had already fixed the problem. When a domain starts serving the crawler again, its next scan removes it from here without anyone editing this page.

DomainCrawlerrobots.txtWhat the edge returned
wikipedia.orgClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
yahoo.comOAI-SearchBotpermits in robots.txtrefused at the edge, measured 2026-08-23
roblox.comOAI-SearchBotpermits in robots.txtrefused at the edge, measured 2026-08-23
cloudflare.netGPTBotpermits in robots.txtrefused at the edge, measured 2026-08-23
wordpress.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
nih.govClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
forms.gleClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
xiaomi.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
tumblr.comGPTBotpermits in robots.txtrefused at the edge, measured 2026-08-23
rubiconproject.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
wikimedia.orgClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
ibm.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
springer.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
duckdns.orgClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
wiley.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
salesforce.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
weebly.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
mi.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
gitlab.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
expireddomains.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
espn.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
wp.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
substack.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
att.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
ietf.orgGPTBotpermits in robots.txtrefused at the edge, measured 2026-08-23
cookielaw.orgOAI-SearchBotpermits in robots.txtrefused at the edge, measured 2026-08-23
oup.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
trendmicro.comClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
noaa.govClaudeBotpermits in robots.txtrefused at the edge, measured 2026-08-23
shopee.com.brGPTBotpermits in robots.txtrefused at the edge, measured 2026-08-23

A spot re-probe of 31 of these from an unrelated network agreed with the stored measurement 27 times; 4 had changed since their last scan. Edge behaviour moves, so anything acted on should be re-probed at the moment of acting rather than taken from this page. How this is measured.

Why this is worth a company's attention the two halves disagree

Plenty of sites block AI crawlers deliberately, and that is a legitimate business decision this census does not argue with. This is the other thing: the site's own published policy says the crawler may fetch it, and the infrastructure in front of the site refuses anyway. The two statements cannot both be intentional.

The shape of the refusal usually says which. A bare 403 to one named user agent, where a browser gets 200, is a bot-management rule. A 406 or a 404 returned only to the crawler is a WAF signature matching on user agent - nobody writes a policy that says "tell OpenAI this page does not exist". A managed challenge is a CDN preset that was never scoped to exclude the crawlers the marketing team wants reading the site.

The cost is asymmetric and quiet. The company keeps publishing a robots.txt that invites the crawler, sees no error in its own logs worth investigating, and simply does not appear in answers built from that crawler's index.

Check your own domain free, no account

The full list of 5,144 is public and so is everything behind it.

  • Measure one domain and the report names every crawler permitted in robots.txt and refused at the edge.
  • https://crawlcensus.com/agents/gptbot/refused.txt — the whole list for one crawler, one domain per line, with its count and method in the header. Also claudebot, perplexitybot, oai-searchbot.
  • The full corpus as JSON, 28,874 domains, CC BY 4.0. Reproduce the number yourself rather than taking it from this page.
  • Reproduce a single row with two commands: curl -A "GPTBot/1.2" https://example.com/ against curl https://example.com/.