Open dataset · CC BY 4.0

The AI crawler census

Every domain here was measured the same way: read its robots.txt, resolve the rules for 48 tracked AI agents, then send live requests as four of them and record what came back. No survey, no self-reporting. Figures update as the crawler works through the queue.

7,054
domains measured
70.1
mean AI access score
9%
block an answer engine
19%
refuse AI agents at the edge
6%
of training-crawler rules are blocks
2%
of answer-index rules are blocks
3%
of live-retrieval rules are blocks
13%
publish an llms.txt

How the web's posture is moving building the series

The daily series starts once the corpus has been measured on two separate days. Come back tomorrow.

Block rate by crawler the headline table

CrawlerOperatorContent used forBlock rateSites blockingSites measured
CCBot Common Crawl Foundation archive 13.3% 315 2,373
GPTBot OpenAI training 12.8% 304 2,373
Bytespider ByteDance training 12.2% 290 2,373
ClaudeBot Anthropic training 11.5% 272 2,373
Meta-ExternalAgent Meta training 10.6% 251 2,373
Google-Extended Google training 10.3% 245 2,373
Applebot-Extended Apple training 9.3% 220 2,373
Amazonbot Amazon training 8.7% 206 2,373
omgilibot Webz.io training 7.4% 176 2,373
Diffbot Diffbot training 7.4% 175 2,373
cohere-ai Cohere live retrieval 7.1% 169 2,373
omgili Webz.io training 7.0% 166 2,373
PerplexityBot Perplexity answer index 6.9% 163 2,373
anthropic-ai Anthropic training 6.4% 151 2,373
ChatGPT-User OpenAI live retrieval 5.9% 140 2,373
Timpibot Timpi training 5.6% 132 2,373
FacebookBot Meta training 5.5% 131 2,373
YouBot You.com answer index 5.4% 128 2,373
ImagesiftBot ImageSift (Hive) training 5.0% 119 2,373
Scrapy Zyte (open-source framework) training 4.6% 108 2,373
Meta-ExternalFetcher Meta live retrieval 4.1% 98 2,373
AI2Bot Allen Institute for AI training 4.1% 97 2,373
OAI-SearchBot OpenAI answer index 4.0% 95 2,373
cohere-training-data-crawler Cohere training 3.9% 93 2,373
Claude-User Anthropic live retrieval 3.9% 92 2,373
DuckAssistBot DuckDuckGo live retrieval 3.9% 92 2,373
PanguBot Huawei training 3.8% 90 2,373
Claude-SearchBot Anthropic answer index 3.8% 89 2,373
Perplexity-User Perplexity live retrieval 3.5% 84 2,373
Kangaroo Bot Kangaroo LLM training 3.5% 82 2,373
MistralAI-User Mistral AI live retrieval 3.4% 80 2,373
SemrushBot-OCOB Semrush training 3.2% 76 2,373
Google-CloudVertexBot Google live retrieval 3.2% 75 2,373
GoogleOther Google training 2.4% 56 2,373
Meta-WebIndexer Meta answer index 2.2% 53 2,373
Applebot Apple answer index 1.2% 29 2,373
TikTokSpider ByteDance live retrieval 1.2% 29 2,373
Amzn-SearchBot Amazon answer index 0.8% 20 2,373
Amzn-User Amazon live retrieval 0.7% 17 2,373
facebookexternalhit Meta live retrieval 0.4% 9 2,373
MistralAI-Index Mistral AI answer index 0.4% 9 2,373
MistralAI-Training Mistral AI training 0.3% 6 2,373
Bingbot Microsoft answer index 0.2% 5 2,373
Googlebot-News Google answer index 0.1% 3 2,373
Googlebot Google answer index 0.1% 2 2,373
msnbot Microsoft answer index 0.1% 2 2,373
Meta-ExternalAds Meta training 0.0% 1 2,373
OAI-AdsBot OpenAI live retrieval 0.0% 1 2,373

Block rate is the share of measured sites whose robots.txt disallows that agent at the site root. A site with no robots.txt counts as allowing everything, which is what the standard specifies.

Grade distribution

7,054Grade B: 2.64k (37.4%)Grade C: 2.25k (31.8%)Grade A: 1.1k (15.6%)Grade D: 954 (13.5%)Grade F: 88 (1.2%)Grade A+: 35 (0.5%)7,054sites

Blocking by edge provider

EdgeSitesRefuse AI agents
Cloudflare7,05419%

Edge is inferred from response headers. A high rate here usually reflects a managed bot rule set rather than a deliberate choice by the site owner.

Take the data free, attributed

The corpus is published under CC BY 4.0. Attribute as Source: Crawl Census (crawlcensus.com).

Corpus snapshot (JSON) Census totals Crawler registry Change stream Methodology