Top-level domain segment

AI crawler access on .org domains

411 of the 1,678 domains measured in this segment refuse a request from an identified AI user agent at the edge, 24% of the segment, against 6% whose robots.txt disallows an answer engine at the site root. The block that costs these sites citations is not the one written in the file. That answer-engine rate is 1.1 points below the corpus rate of 7%, which is itself measured across 30,319 domains.

Mean AI access score across those 1,678 domains is 72.7 of 100, 147 of them publish an llms.txt (9%), and 645 publish structured data an answer engine can parse (38%).

1,678
domains measured
72.7
mean AI access score of 1,678
6%
block an answer engine (93 of 1,678)
9%
publish llms.txt (147 of 1,678)

Block rate by crawler, inside this segment robots.txt at the site root

CrawlerOperatorContent used forBlock rateBlockingMeasured
GPTBot OpenAI training 6.5% 109 1,677
Bytespider ByteDance training 5.5% 93 1,677
CCBot Common Crawl Foundation archive 5.4% 90 1,677
ClaudeBot Anthropic training 4.5% 75 1,677
Google-Extended Google training 4.2% 70 1,677
Amazonbot Amazon training 3.9% 66 1,677
SemrushBot-OCOB Semrush training 3.9% 65 1,677
Applebot-Extended Apple training 3.6% 61 1,677
Meta-ExternalAgent Meta training 3.2% 54 1,677
anthropic-ai Anthropic training 3.2% 53 1,677
omgilibot Webz.io training 3.0% 50 1,677
omgili Webz.io training 2.9% 49 1,677
Diffbot Diffbot training 2.8% 47 1,677
cohere-ai Cohere live retrieval 2.8% 47 1,677
ChatGPT-User OpenAI live retrieval 2.7% 45 1,677
PerplexityBot Perplexity answer index 2.4% 41 1,677
FacebookBot Meta training 2.4% 41 1,677
YouBot You.com answer index 1.7% 29 1,677
ImagesiftBot ImageSift (Hive) training 1.7% 29 1,677
GoogleOther Google training 1.5% 25 1,677
Scrapy Zyte (open-source framework) training 1.4% 24 1,677
Applebot Apple answer index 1.4% 23 1,677
Timpibot Timpi training 1.1% 19 1,677
OAI-SearchBot OpenAI answer index 1.0% 17 1,677
Google-CloudVertexBot Google live retrieval 1.0% 17 1,677
AI2Bot Allen Institute for AI training 1.0% 16 1,677
Meta-ExternalFetcher Meta live retrieval 0.9% 15 1,677
cohere-training-data-crawler Cohere training 0.9% 15 1,677
TikTokSpider ByteDance live retrieval 0.7% 12 1,677
PanguBot Huawei training 0.7% 12 1,677
Kangaroo Bot Kangaroo LLM training 0.6% 10 1,677
DuckAssistBot DuckDuckGo live retrieval 0.6% 10 1,677
MistralAI-Training Mistral AI training 0.5% 9 1,677
Claude-User Anthropic live retrieval 0.5% 9 1,677
Claude-SearchBot Anthropic answer index 0.5% 9 1,677
Perplexity-User Perplexity live retrieval 0.5% 8 1,677
MistralAI-User Mistral AI live retrieval 0.4% 7 1,677
Meta-WebIndexer Meta answer index 0.2% 4 1,677
facebookexternalhit Meta live retrieval 0.2% 4 1,677
msnbot Microsoft answer index 0.2% 3 1,677
Bingbot Microsoft answer index 0.1% 2 1,677
Amzn-User Amazon live retrieval 0.1% 2 1,677
Amzn-SearchBot Amazon answer index 0.1% 2 1,677
MistralAI-Index Mistral AI answer index 0.1% 1 1,677
Meta-ExternalAds Meta training 0.1% 1 1,677
OAI-AdsBot OpenAI live retrieval 0.0% 0 1,677
Googlebot-News Google answer index 0.0% 0 1,677
Googlebot Google answer index 0.0% 0 1,677

Block rate is the share of this segment's domains holding a policy record for that agent whose robots.txt disallows it at the site root, shown with its numerator and denominator. Denominators differ between agents because a record only exists once a domain has been measured for it. A domain serving no robots.txt counts as allowing everything, which is what the standard specifies.

Most accessible

DomainScoreEngines blockedllms.txt
1 worldhistory.org 99 0 yes
2 getgrav.org 99 0 yes
3 23andme.org 98 0 yes
4 taxfoundation.org 98 0 yes
5 wri.org 98 0 yes
6 justsecurity.org 97 0 yes
7 findmykids.org 97 0 yes
8 ccl.org 97 0 yes
9 skincancer.org 97 0 yes
10 nextjs.org 97 0 yes
11 visionofhumanity.org 97 0 yes
12 commondreams.org 96 0 yes
13 understood.org 96 0 yes
14 esrb.org 96 0 yes
15 responsiblestatecraft.org 96 0 yes

Highest AI access scores among the 1,678 domains measured in this segment.

Least accessible

DomainScoreEngines blockedllms.txt
1 stress.org 25 2 no
2 unv.org 27 0 no
3 pcre.org 28 0 no
4 transformativeworks.org 28 0 no
5 americanpregnancy.org 28 0 no
6 independent.org 28 0 no
7 oxfordjournals.org 29 0 no
8 artuk.org 30 0 no
9 cato.org 31 0 no
10 aappublications.org 32 0 no
11 scitation.org 32 0 no
12 lightpokies.org 32 0 no
13 mystonks.org 33 0 no
14 perchance.org 34 0 no
15 metopera.org 34 0 no

Lowest AI access scores among the same 1,678 domains.

Compare with the rest of the grouping top-level domain segments

The largest 12 other segments in this grouping, ordered by size. The share beside each is that segment's own answer-engine block rate over its own count. Full table on the segments index.

What this segment does not tell you read this before citing it

A top-level domain is a weak proxy for jurisdiction. Most of them are open to any registrant anywhere, so a domain under a country-code registry may be operated from outside that country by a company no local law reaches. Read these rows as differences between registration communities and the publishing cultures inside them, not as differences between legal regimes.

Measure your own domain

The audit that produced every number on this page runs on any domain in about six seconds: robots.txt resolved for 48 agents, then live requests sent as four of them to see whether the edge agrees with the file.