Top-level domain segment

AI crawler access on .uk domains

55 of the 202 domains measured in this segment refuse a request from an identified AI user agent at the edge, 27% of the segment, against 16% whose robots.txt disallows an answer engine at the site root. The block that costs these sites citations is not the one written in the file. That answer-engine rate is 2.2 times the corpus rate of 7%, which is itself measured across 16,934 domains.

Mean AI access score across those 202 domains is 71.4 of 100, 14 of them publish an llms.txt (7%), and 99 publish structured data an answer engine can parse (49%).

202
domains measured
71.4
mean AI access score of 202
16%
block an answer engine (32 of 202)
7%
publish llms.txt (14 of 202)

Block rate by crawler, inside this segment robots.txt at the site root

CrawlerOperatorContent used forBlock rateBlockingMeasured
CCBot Common Crawl Foundation archive 19.8% 40 202
GPTBot OpenAI training 18.8% 38 202
ClaudeBot Anthropic training 18.3% 37 202
Meta-ExternalAgent Meta training 16.3% 33 202
Applebot-Extended Apple training 16.3% 33 202
Bytespider ByteDance training 14.9% 30 202
Google-Extended Google training 14.4% 29 202
PerplexityBot Perplexity answer index 13.4% 27 202
anthropic-ai Anthropic training 12.4% 25 202
YouBot You.com answer index 11.4% 23 202
Amazonbot Amazon training 11.4% 23 202
OAI-SearchBot OpenAI answer index 9.4% 19 202
Diffbot Diffbot training 7.9% 16 202
omgilibot Webz.io training 6.9% 14 202
cohere-ai Cohere live retrieval 6.9% 14 202
Timpibot Timpi training 6.4% 13 202
Perplexity-User Perplexity live retrieval 6.4% 13 202
omgili Webz.io training 6.4% 13 202
Google-CloudVertexBot Google live retrieval 6.4% 13 202
ImagesiftBot ImageSift (Hive) training 5.9% 12 202
FacebookBot Meta training 5.9% 12 202
DuckAssistBot DuckDuckGo live retrieval 5.9% 12 202
Claude-SearchBot Anthropic answer index 5.9% 12 202
Meta-ExternalFetcher Meta live retrieval 5.4% 11 202
Claude-User Anthropic live retrieval 5.4% 11 202
ChatGPT-User OpenAI live retrieval 5.4% 11 202
GoogleOther Google training 5.0% 10 202
Scrapy Zyte (open-source framework) training 4.5% 9 202
PanguBot Huawei training 4.5% 9 202
MistralAI-User Mistral AI live retrieval 4.5% 9 202
SemrushBot-OCOB Semrush training 4.0% 8 202
Kangaroo Bot Kangaroo LLM training 3.0% 6 202
cohere-training-data-crawler Cohere training 3.0% 6 202
AI2Bot Allen Institute for AI training 3.0% 6 202
Meta-WebIndexer Meta answer index 2.5% 5 202
TikTokSpider ByteDance live retrieval 2.0% 4 202
Amzn-SearchBot Amazon answer index 1.5% 3 202
Amzn-User Amazon live retrieval 1.0% 2 202
Meta-ExternalAds Meta training 0.5% 1 202
Applebot Apple answer index 0.5% 1 202
OAI-AdsBot OpenAI live retrieval 0.0% 0 202
msnbot Microsoft answer index 0.0% 0 202
MistralAI-Training Mistral AI training 0.0% 0 202
MistralAI-Index Mistral AI answer index 0.0% 0 202
Googlebot-News Google answer index 0.0% 0 202
Googlebot Google answer index 0.0% 0 202
facebookexternalhit Meta live retrieval 0.0% 0 202
Bingbot Microsoft answer index 0.0% 0 202

Block rate is the share of this segment's domains holding a policy record for that agent whose robots.txt disallows it at the site root, shown with its numerator and denominator. Denominators differ between agents because a record only exists once a domain has been measured for it. A domain serving no robots.txt counts as allowing everything, which is what the standard specifies.

Most accessible

DomainScoreEngines blockedllms.txt
1 lrb.co.uk 91 1 no
2 reed.co.uk 91 0 yes
3 vivastreet.co.uk 91 0 yes
4 ebi.ac.uk 91 0 no
5 screamingfrog.co.uk 90 0 no
6 fca.org.uk 90 0 no
7 kcl.ac.uk 90 0 no
8 spectator.co.uk 88 0 no
9 matalan.co.uk 87 0 no
10 ageuk.org.uk 87 0 no
11 fasthosts.co.uk 87 0 yes
12 kent.ac.uk 87 0 no
13 wwf.org.uk 87 0 no
14 port.ac.uk 87 0 yes
15 zen.co.uk 86 0 no

Highest AI access scores among the 202 domains measured in this segment.

Least accessible

DomainScoreEngines blockedllms.txt
1 amazon.co.uk 33 13 no
2 twinkl.co.uk 33 4 no
3 mirror.co.uk 37 3 no
4 manchestereveningnews.co.uk 37 3 no
5 walesonline.co.uk 37 3 no
6 dailystar.co.uk 37 3 no
7 liverpoolecho.co.uk 37 3 no
8 dailyrecord.co.uk 37 3 no
9 birminghammail.co.uk 37 3 no
10 chroniclelive.co.uk 37 3 no
11 bristolpost.co.uk 37 3 no
12 49s.co.uk 41 0 no
13 express.co.uk 44 3 no
14 espn.co.uk 44 1 no
15 guim.co.uk 47 0 no

Lowest AI access scores among the same 202 domains.

Compare with the rest of the grouping top-level domain segments

The largest 12 other segments in this grouping, ordered by size. The share beside each is that segment's own answer-engine block rate over its own count. Full table on the segments index.

What this segment does not tell you read this before citing it

A top-level domain is a weak proxy for jurisdiction. Most of them are open to any registrant anywhere, so a domain under a country-code registry may be operated from outside that country by a company no local law reaches. Read these rows as differences between registration communities and the publishing cultures inside them, not as differences between legal regimes.

Measure your own domain

The audit that produced every number on this page runs on any domain in about six seconds: robots.txt resolved for 48 agents, then live requests sent as four of them to see whether the edge agrees with the file.