Open dataset · CC BY 4.0

The AI crawler census

Every domain here was measured the same way: read its robots.txt, resolve the rules for 48 tracked AI agents, then send live requests as four of them and record what came back. No survey, no self-reporting. Figures update as the crawler works through the queue.

What share of the web can an AI crawler actually read the honest denominator

The headline rate elsewhere on this site, 19%, is measured across the 30,185 websites that served this crawler a page. It leaves out the 5,739 websites that refused before any question about AI could be asked - a bare 403 to an ordinary browser user agent, no robots.txt consulted, no content returned.

Excluding them understates the answer, because a site that refuses every programmatic client refuses AI crawlers most of all. Counted properly, of 35,924 websites in the corpus, 11,560 - 32% - cannot be read by an AI crawler: 5,739 refuse outright and 5,821 of the remainder refuse an AI user agent specifically.

19%
refuse AI agents, of sites that answered
32%
unreadable by an AI crawler, of all websites
5,739
refuse every programmatic client

The same exclusion could have distorted the robots.txt figures, and it does not. Those 5,739 sites refused to serve a page but still served a policy file, because robots.txt is fetched in the same pass rather than after the home page succeeds. Their stated policy is indistinguishable from everyone else's: 6.7% disallow an answer engine against 6.6% among sites that answered, and 11.4% disallow a training crawler against 11.7%. A site's written policy turns out to be independent of whether its infrastructure tolerates a programmatic client at all - which is the same split this census keeps finding between what operators decide and what their stack does by default. The robots.txt rates elsewhere on this site are therefore left on the narrower denominator, because widening it changes nothing.

Both numbers are published because they answer different questions. The narrower one is what this crawler directly measured. The wider one is what a person asking "how much of the web can a model actually read" should be told, and it is nearly twice as large. Among the sites refusing outright are openai.com, chatgpt.com and perplexity.ai, whose own products depend on crawling everyone else.

What the corpus is made of 49,557 hostnames

The seed list ranks hostnames by traffic, and a large share of the most trafficked hostnames on the internet are not websites: CDN endpoints, update services and API hosts that answer enormous numbers of DNS queries and serve no page at their apex. akamaiedge.net, cloudfront.net, windowsupdate.com and aaplimg.com all sit in the first hundred and none of them has a home page.

OutcomeHostnamesShare
Website measured30,18560.9%
No website at this hostname12,64925.5%
Refused before assessment5,73911.6%
Other HTTP error9842.0%

Every rate published on this site divides by the first row only, because a hostname with no website cannot have a crawler policy. The distinction was verified rather than assumed: 32 hostnames returning a Cloudflare origin error were fetched again from a residential connection with redirects followed, and none of them served a page there either.

How much of this depends on how prominent a site is the corpus is Tranco's top 36,000

Every figure on this page describes domains drawn from the Tranco popularity ranking, positions 2 to 35,634. It is not a sample of the web, which is mostly small and obscure, and rates measured here should not be read as web-wide. The gradient below is published so the direction and size of that bias is visible rather than merely disclosed.

Popularity bandDomainsRefuse at the edgeBlock an answer enginePublish llms.txtMean score
Top 1,00048413.8%12%17.8%72.2
1,000 to 10,0005,19120%9%14.3%73.3
10,000 and below24,40419.3%6%13.4%73.4
Unranked10614.2%0.9%31.1%78.1

The two block types behave differently, and that is the finding. Declaring a block in robots.txt tracks prominence closely: the most visited thousand domains do it at roughly twice the rate of those ranked ten thousand and below. Refusing an AI user agent at the edge does not follow that curve at all - it peaks in the middle of the ranking, among sites large enough to sit behind a managed WAF but not large enough to have anyone tuning its bot rules. The first number is a decision. The second is mostly a default.

The measurement

Across 30,185 domains measured by direct request, 7% disallow at least one answer engine in robots.txt and 12% disallow at least one training crawler. Independently of robots.txt, 19% refuse an AI crawler user agent at the network edge. 14% publish an llms.txt.

Measured 2026-10-06 by direct request to every domain in the corpus. Method: how this is measured. Reuse under CC BY 4.0 with attribution.

Cite as: Crawl Census, "The AI crawler census", measured 2026-10-06. https://crawlcensus.com/census

30,185
domains measured
73.4
mean AI access score
7%
block an answer engine
19%
refuse AI agents at the edge
3%
of training-crawler rules are blocks
1%
of answer-index rules are blocks
2%
of live-retrieval rules are blocks
14%
publish an llms.txt

How the web's posture is moving share of measured sites, by day

% of sites20151050% of sites08-2209-1310-06block an answer engineblock a training crawlerpublish llms.txt

Block rate by crawler the headline table

CrawlerOperatorContent used forBlock rateSites blockingSites measured
GPTBot OpenAI training 7.1% 2,145 30,177
CCBot Common Crawl Foundation archive 6.8% 2,066 30,177
Bytespider ByteDance training 6.5% 1,950 30,177
ClaudeBot Anthropic training 5.8% 1,743 30,177
Meta-ExternalAgent Meta training 4.7% 1,430 30,177
Amazonbot Amazon training 4.7% 1,411 30,177
Google-Extended Google training 4.6% 1,391 30,177
Applebot-Extended Apple training 4.6% 1,378 30,177
anthropic-ai Anthropic training 4.4% 1,320 30,177
omgilibot Webz.io training 4.3% 1,289 30,177
Diffbot Diffbot training 4.1% 1,224 30,177
cohere-ai Cohere live retrieval 4.1% 1,223 30,177
omgili Webz.io training 3.9% 1,171 30,177
SemrushBot-OCOB Semrush training 3.7% 1,126 30,177
PerplexityBot Perplexity answer index 3.5% 1,054 30,177
ChatGPT-User OpenAI live retrieval 3.5% 1,053 30,177
FacebookBot Meta training 3.2% 979 30,177
YouBot You.com answer index 2.9% 890 30,177
ImagesiftBot ImageSift (Hive) training 2.8% 853 30,177
Timpibot Timpi training 2.8% 833 30,177
Scrapy Zyte (open-source framework) training 2.5% 745 30,177
AI2Bot Allen Institute for AI training 2.3% 706 30,177
OAI-SearchBot OpenAI answer index 2.2% 650 30,177
cohere-training-data-crawler Cohere training 2.1% 620 30,177
Meta-ExternalFetcher Meta live retrieval 1.8% 557 30,177
DuckAssistBot DuckDuckGo live retrieval 1.7% 519 30,177
PanguBot Huawei training 1.7% 518 30,177
Claude-User Anthropic live retrieval 1.7% 498 30,177
Claude-SearchBot Anthropic answer index 1.6% 484 30,177
Perplexity-User Perplexity live retrieval 1.6% 469 30,177
Kangaroo Bot Kangaroo LLM training 1.5% 466 30,177
MistralAI-User Mistral AI live retrieval 1.5% 447 30,177
Google-CloudVertexBot Google live retrieval 1.4% 423 30,177
GoogleOther Google training 1.3% 383 30,177
Meta-WebIndexer Meta answer index 1.2% 359 30,177
Applebot Apple answer index 1.1% 343 30,177
TikTokSpider ByteDance live retrieval 0.8% 255 30,177
Amzn-SearchBot Amazon answer index 0.8% 233 30,177
Amzn-User Amazon live retrieval 0.6% 167 30,177
MistralAI-Training Mistral AI training 0.5% 142 30,177
MistralAI-Index Mistral AI answer index 0.3% 81 30,177
facebookexternalhit Meta live retrieval 0.2% 53 30,177
Bingbot Microsoft answer index 0.1% 40 30,177
Googlebot-News Google answer index 0.1% 36 30,177
Meta-ExternalAds Meta training 0.1% 22 30,177
msnbot Microsoft answer index 0.1% 21 30,177
Googlebot Google answer index 0.1% 16 30,177
OAI-AdsBot OpenAI live retrieval 0.0% 11 30,177

Block rate is the share of measured sites whose robots.txt disallows that agent at the site root. A site with no robots.txt counts as allowing everything, which is what the standard specifies.

Why this table's total is lower than 30,185. The headline counts every domain that answered over HTTPS. This table counts only domains that already have a stored policy record for that specific agent, which is a snapshot taken at the last daily rollup and therefore trails the live crawl. The two converge as the queue drains. Both numbers are shown rather than reconciled silently.

Grade distribution

30,185Grade B: 10.73k (35.5%)Grade A: 7.43k (24.6%)Grade C: 7.31k (24.2%)Grade D: 3.32k (11%)Grade A+: 1.03k (3.4%)Grade F: 374 (1.2%)30,185sites

Blocking by top-level domain

TLDSitesBlock an answer enginePublish llms.txt
.com14,7817%17%
.org1,6666%9%
.net1,3014%10%
.ru1,1653%9%
.de82516%7%
.io6612%29%
.edu4712%7%
.uk44717%9%
.jp4169%5%
.br3746%13%
.fr32915%3%
.in3043%10%

A TLD is a weak proxy for jurisdiction, since most are open to anyone, but the spread between them is large and consistent. Full segmentation.

Take the data free, attributed

The corpus is published under CC BY 4.0. Attribute as Source: Crawl Census (crawlcensus.com).

Corpus snapshot (JSON) Census totals Crawler registry Change stream Methodology