AI crawler access on .uk domains
55 of the 202 domains measured in this segment refuse a request from an identified AI user agent at the edge, 27% of the segment, against 16% whose robots.txt disallows an answer engine at the site root. The block that costs these sites citations is not the one written in the file. That answer-engine rate is 2.2 times the corpus rate of 7%, which is itself measured across 16,934 domains.
Mean AI access score across those 202 domains is 71.4 of 100, 14 of them publish an llms.txt (7%), and 99 publish structured data an answer engine can parse (49%).
Block rate by crawler, inside this segment robots.txt at the site root
| Crawler | Operator | Content used for | Block rate | Blocking | Measured |
|---|---|---|---|---|---|
| CCBot | Common Crawl Foundation | archive | 19.8% | 40 | 202 |
| GPTBot | OpenAI | training | 18.8% | 38 | 202 |
| ClaudeBot | Anthropic | training | 18.3% | 37 | 202 |
| Meta-ExternalAgent | Meta | training | 16.3% | 33 | 202 |
| Applebot-Extended | Apple | training | 16.3% | 33 | 202 |
| Bytespider | ByteDance | training | 14.9% | 30 | 202 |
| Google-Extended | training | 14.4% | 29 | 202 | |
| PerplexityBot | Perplexity | answer index | 13.4% | 27 | 202 |
| anthropic-ai | Anthropic | training | 12.4% | 25 | 202 |
| YouBot | You.com | answer index | 11.4% | 23 | 202 |
| Amazonbot | Amazon | training | 11.4% | 23 | 202 |
| OAI-SearchBot | OpenAI | answer index | 9.4% | 19 | 202 |
| Diffbot | Diffbot | training | 7.9% | 16 | 202 |
| omgilibot | Webz.io | training | 6.9% | 14 | 202 |
| cohere-ai | Cohere | live retrieval | 6.9% | 14 | 202 |
| Timpibot | Timpi | training | 6.4% | 13 | 202 |
| Perplexity-User | Perplexity | live retrieval | 6.4% | 13 | 202 |
| omgili | Webz.io | training | 6.4% | 13 | 202 |
| Google-CloudVertexBot | live retrieval | 6.4% | 13 | 202 | |
| ImagesiftBot | ImageSift (Hive) | training | 5.9% | 12 | 202 |
| FacebookBot | Meta | training | 5.9% | 12 | 202 |
| DuckAssistBot | DuckDuckGo | live retrieval | 5.9% | 12 | 202 |
| Claude-SearchBot | Anthropic | answer index | 5.9% | 12 | 202 |
| Meta-ExternalFetcher | Meta | live retrieval | 5.4% | 11 | 202 |
| Claude-User | Anthropic | live retrieval | 5.4% | 11 | 202 |
| ChatGPT-User | OpenAI | live retrieval | 5.4% | 11 | 202 |
| GoogleOther | training | 5.0% | 10 | 202 | |
| Scrapy | Zyte (open-source framework) | training | 4.5% | 9 | 202 |
| PanguBot | Huawei | training | 4.5% | 9 | 202 |
| MistralAI-User | Mistral AI | live retrieval | 4.5% | 9 | 202 |
| SemrushBot-OCOB | Semrush | training | 4.0% | 8 | 202 |
| Kangaroo Bot | Kangaroo LLM | training | 3.0% | 6 | 202 |
| cohere-training-data-crawler | Cohere | training | 3.0% | 6 | 202 |
| AI2Bot | Allen Institute for AI | training | 3.0% | 6 | 202 |
| Meta-WebIndexer | Meta | answer index | 2.5% | 5 | 202 |
| TikTokSpider | ByteDance | live retrieval | 2.0% | 4 | 202 |
| Amzn-SearchBot | Amazon | answer index | 1.5% | 3 | 202 |
| Amzn-User | Amazon | live retrieval | 1.0% | 2 | 202 |
| Meta-ExternalAds | Meta | training | 0.5% | 1 | 202 |
| Applebot | Apple | answer index | 0.5% | 1 | 202 |
| OAI-AdsBot | OpenAI | live retrieval | 0.0% | 0 | 202 |
| msnbot | Microsoft | answer index | 0.0% | 0 | 202 |
| MistralAI-Training | Mistral AI | training | 0.0% | 0 | 202 |
| MistralAI-Index | Mistral AI | answer index | 0.0% | 0 | 202 |
| Googlebot-News | answer index | 0.0% | 0 | 202 | |
| Googlebot | answer index | 0.0% | 0 | 202 | |
| facebookexternalhit | Meta | live retrieval | 0.0% | 0 | 202 |
| Bingbot | Microsoft | answer index | 0.0% | 0 | 202 |
Block rate is the share of this segment's domains holding a policy record for that agent whose robots.txt disallows it at the site root, shown with its numerator and denominator. Denominators differ between agents because a record only exists once a domain has been measured for it. A domain serving no robots.txt counts as allowing everything, which is what the standard specifies.
Most accessible
| Domain | Score | Engines blocked | llms.txt | |
|---|---|---|---|---|
| 1 | lrb.co.uk | 91 | 1 | no |
| 2 | reed.co.uk | 91 | 0 | yes |
| 3 | vivastreet.co.uk | 91 | 0 | yes |
| 4 | ebi.ac.uk | 91 | 0 | no |
| 5 | screamingfrog.co.uk | 90 | 0 | no |
| 6 | fca.org.uk | 90 | 0 | no |
| 7 | kcl.ac.uk | 90 | 0 | no |
| 8 | spectator.co.uk | 88 | 0 | no |
| 9 | matalan.co.uk | 87 | 0 | no |
| 10 | ageuk.org.uk | 87 | 0 | no |
| 11 | fasthosts.co.uk | 87 | 0 | yes |
| 12 | kent.ac.uk | 87 | 0 | no |
| 13 | wwf.org.uk | 87 | 0 | no |
| 14 | port.ac.uk | 87 | 0 | yes |
| 15 | zen.co.uk | 86 | 0 | no |
Highest AI access scores among the 202 domains measured in this segment.
Least accessible
| Domain | Score | Engines blocked | llms.txt | |
|---|---|---|---|---|
| 1 | amazon.co.uk | 33 | 13 | no |
| 2 | twinkl.co.uk | 33 | 4 | no |
| 3 | mirror.co.uk | 37 | 3 | no |
| 4 | manchestereveningnews.co.uk | 37 | 3 | no |
| 5 | walesonline.co.uk | 37 | 3 | no |
| 6 | dailystar.co.uk | 37 | 3 | no |
| 7 | liverpoolecho.co.uk | 37 | 3 | no |
| 8 | dailyrecord.co.uk | 37 | 3 | no |
| 9 | birminghammail.co.uk | 37 | 3 | no |
| 10 | chroniclelive.co.uk | 37 | 3 | no |
| 11 | bristolpost.co.uk | 37 | 3 | no |
| 12 | 49s.co.uk | 41 | 0 | no |
| 13 | express.co.uk | 44 | 3 | no |
| 14 | espn.co.uk | 44 | 1 | no |
| 15 | guim.co.uk | 47 | 0 | no |
Lowest AI access scores among the same 202 domains.
Compare with the rest of the grouping top-level domain segments
The largest 12 other segments in this grouping, ordered by size. The share beside each is that segment's own answer-engine block rate over its own count. Full table on the segments index.
What this segment does not tell you read this before citing it
A top-level domain is a weak proxy for jurisdiction. Most of them are open to any registrant anywhere, so a domain under a country-code registry may be operated from outside that country by a company no local law reaches. Read these rows as differences between registration communities and the publishing cultures inside them, not as differences between legal regimes.
Measure your own domain
The audit that produced every number on this page runs on any domain in about six seconds: robots.txt resolved for 48 agents, then live requests sent as four of them to see whether the edge agrees with the file.