Top-level domain segment

AI crawler access on .news domains

3 of the 10 domains measured in this segment refuse a request from an identified AI user agent at the edge, 30% of the segment, against 20% whose robots.txt disallows an answer engine at the site root. The block that costs these sites citations is not the one written in the file. That answer-engine rate is three times the corpus rate of 7%, which is itself measured across 20,812 domains.

Mean AI access score across those 10 domains is 62.7 of 100, 0 of them publish an llms.txt (0%), and 3 publish structured data an answer engine can parse (30%). At 10 domains this segment is small enough that one site changing its robots.txt moves every rate on this page.

10
domains measured
62.7
mean AI access score of 10
20%
block an answer engine (2 of 10)
0%
publish llms.txt (0 of 10)

Block rate by crawler, inside this segment robots.txt at the site root

CrawlerOperatorContent used forBlock rateBlockingMeasured
Meta-ExternalAgent Meta training 30.0% 3 10
GPTBot OpenAI training 30.0% 3 10
ClaudeBot Anthropic training 30.0% 3 10
CCBot Common Crawl Foundation archive 30.0% 3 10
Applebot-Extended Apple training 30.0% 3 10
PerplexityBot Perplexity answer index 20.0% 2 10
OAI-SearchBot OpenAI answer index 20.0% 2 10
Google-Extended Google training 20.0% 2 10
Bytespider ByteDance training 20.0% 2 10
anthropic-ai Anthropic training 20.0% 2 10
Amazonbot Amazon training 20.0% 2 10
YouBot You.com answer index 10.0% 1 10
Timpibot Timpi training 10.0% 1 10
Scrapy Zyte (open-source framework) training 10.0% 1 10
Perplexity-User Perplexity live retrieval 10.0% 1 10
omgilibot Webz.io training 10.0% 1 10
omgili Webz.io training 10.0% 1 10
Meta-ExternalFetcher Meta live retrieval 10.0% 1 10
FacebookBot Meta training 10.0% 1 10
Diffbot Diffbot training 10.0% 1 10
cohere-training-data-crawler Cohere training 10.0% 1 10
cohere-ai Cohere live retrieval 10.0% 1 10
Claude-User Anthropic live retrieval 10.0% 1 10
Claude-SearchBot Anthropic answer index 10.0% 1 10
ChatGPT-User OpenAI live retrieval 10.0% 1 10
TikTokSpider ByteDance live retrieval 0.0% 0 10
SemrushBot-OCOB Semrush training 0.0% 0 10
PanguBot Huawei training 0.0% 0 10
OAI-AdsBot OpenAI live retrieval 0.0% 0 10
msnbot Microsoft answer index 0.0% 0 10
MistralAI-User Mistral AI live retrieval 0.0% 0 10
MistralAI-Training Mistral AI training 0.0% 0 10
MistralAI-Index Mistral AI answer index 0.0% 0 10
Meta-WebIndexer Meta answer index 0.0% 0 10
Meta-ExternalAds Meta training 0.0% 0 10
Kangaroo Bot Kangaroo LLM training 0.0% 0 10
ImagesiftBot ImageSift (Hive) training 0.0% 0 10
GoogleOther Google training 0.0% 0 10
Googlebot-News Google answer index 0.0% 0 10
Googlebot Google answer index 0.0% 0 10
Google-CloudVertexBot Google live retrieval 0.0% 0 10
facebookexternalhit Meta live retrieval 0.0% 0 10
DuckAssistBot DuckDuckGo live retrieval 0.0% 0 10
Bingbot Microsoft answer index 0.0% 0 10
Applebot Apple answer index 0.0% 0 10
Amzn-User Amazon live retrieval 0.0% 0 10
Amzn-SearchBot Amazon answer index 0.0% 0 10
AI2Bot Allen Institute for AI training 0.0% 0 10

Block rate is the share of this segment's domains holding a policy record for that agent whose robots.txt disallows it at the site root, shown with its numerator and denominator. Denominators differ between agents because a record only exists once a domain has been measured for it. A domain serving no robots.txt counts as allowing everything, which is what the standard specifies.

Most accessible

DomainScoreEngines blockedllms.txt
1 3isk.news 88 0 no
2 jalalive.news 86 0 no
3 hochi.news 70 8 no
4 elbalad.news 68 0 no
5 ground.news 64 0 no
6 bad.news 59 0 no
7 elif.news 55 0 no
8 apple.news 48 0 no
9 psiphon.news 46 0 no
10 mylondon.news 43 3 no

Highest AI access scores among the 10 domains measured in this segment.

Least accessible

DomainScoreEngines blockedllms.txt
1 mylondon.news 43 3 no
2 psiphon.news 46 0 no
3 apple.news 48 0 no
4 elif.news 55 0 no
5 bad.news 59 0 no
6 ground.news 64 0 no
7 elbalad.news 68 0 no
8 hochi.news 70 8 no
9 jalalive.news 86 0 no
10 3isk.news 88 0 no

Lowest AI access scores among the same 10 domains.

Compare with the rest of the grouping top-level domain segments

The largest 12 other segments in this grouping, ordered by size. The share beside each is that segment's own answer-engine block rate over its own count. Full table on the segments index.

What this segment does not tell you read this before citing it

A top-level domain is a weak proxy for jurisdiction. Most of them are open to any registrant anywhere, so a domain under a country-code registry may be operated from outside that country by a company no local law reaches. Read these rows as differences between registration communities and the publishing cultures inside them, not as differences between legal regimes.

Measure your own domain

The audit that produced every number on this page runs on any domain in about six seconds: robots.txt resolved for 48 agents, then live requests sent as four of them to see whether the edge agrees with the file.