AI crawler access on .dev domains
7 of the 63 domains measured in this segment refuse a request from an identified AI user agent at the edge, 11% of the segment, against 2% whose robots.txt disallows an answer engine at the site root. The block that costs these sites citations is not the one written in the file. That answer-engine rate is 4.9 points below the corpus rate of 7%, which is itself measured across 27,473 domains.
Mean AI access score across those 63 domains is 72.1 of 100, 19 of them publish an llms.txt (30%), and 26 publish structured data an answer engine can parse (41%).
Block rate by crawler, inside this segment robots.txt at the site root
| Crawler | Operator | Content used for | Block rate | Blocking | Measured |
|---|---|---|---|---|---|
| Bytespider | ByteDance | training | 7.9% | 5 | 63 |
| Meta-ExternalAgent | Meta | training | 4.8% | 3 | 63 |
| GPTBot | OpenAI | training | 4.8% | 3 | 63 |
| Google-Extended | training | 4.8% | 3 | 63 | |
| ClaudeBot | Anthropic | training | 4.8% | 3 | 63 |
| CCBot | Common Crawl Foundation | archive | 4.8% | 3 | 63 |
| Applebot-Extended | Apple | training | 4.8% | 3 | 63 |
| Amazonbot | Amazon | training | 4.8% | 3 | 63 |
| SemrushBot-OCOB | Semrush | training | 3.2% | 2 | 63 |
| YouBot | You.com | answer index | 1.6% | 1 | 63 |
| Timpibot | Timpi | training | 1.6% | 1 | 63 |
| TikTokSpider | ByteDance | live retrieval | 1.6% | 1 | 63 |
| Scrapy | Zyte (open-source framework) | training | 1.6% | 1 | 63 |
| PerplexityBot | Perplexity | answer index | 1.6% | 1 | 63 |
| Perplexity-User | Perplexity | live retrieval | 1.6% | 1 | 63 |
| PanguBot | Huawei | training | 1.6% | 1 | 63 |
| omgilibot | Webz.io | training | 1.6% | 1 | 63 |
| omgili | Webz.io | training | 1.6% | 1 | 63 |
| OAI-SearchBot | OpenAI | answer index | 1.6% | 1 | 63 |
| MistralAI-User | Mistral AI | live retrieval | 1.6% | 1 | 63 |
| Meta-WebIndexer | Meta | answer index | 1.6% | 1 | 63 |
| Meta-ExternalFetcher | Meta | live retrieval | 1.6% | 1 | 63 |
| Meta-ExternalAds | Meta | training | 1.6% | 1 | 63 |
| Kangaroo Bot | Kangaroo LLM | training | 1.6% | 1 | 63 |
| ImagesiftBot | ImageSift (Hive) | training | 1.6% | 1 | 63 |
| GoogleOther | training | 1.6% | 1 | 63 | |
| Google-CloudVertexBot | live retrieval | 1.6% | 1 | 63 | |
| facebookexternalhit | Meta | live retrieval | 1.6% | 1 | 63 |
| FacebookBot | Meta | training | 1.6% | 1 | 63 |
| DuckAssistBot | DuckDuckGo | live retrieval | 1.6% | 1 | 63 |
| Diffbot | Diffbot | training | 1.6% | 1 | 63 |
| cohere-training-data-crawler | Cohere | training | 1.6% | 1 | 63 |
| cohere-ai | Cohere | live retrieval | 1.6% | 1 | 63 |
| Claude-User | Anthropic | live retrieval | 1.6% | 1 | 63 |
| Claude-SearchBot | Anthropic | answer index | 1.6% | 1 | 63 |
| ChatGPT-User | OpenAI | live retrieval | 1.6% | 1 | 63 |
| Applebot | Apple | answer index | 1.6% | 1 | 63 |
| anthropic-ai | Anthropic | training | 1.6% | 1 | 63 |
| Amzn-User | Amazon | live retrieval | 1.6% | 1 | 63 |
| Amzn-SearchBot | Amazon | answer index | 1.6% | 1 | 63 |
| AI2Bot | Allen Institute for AI | training | 1.6% | 1 | 63 |
| OAI-AdsBot | OpenAI | live retrieval | 0.0% | 0 | 63 |
| msnbot | Microsoft | answer index | 0.0% | 0 | 63 |
| MistralAI-Training | Mistral AI | training | 0.0% | 0 | 63 |
| MistralAI-Index | Mistral AI | answer index | 0.0% | 0 | 63 |
| Googlebot-News | answer index | 0.0% | 0 | 63 | |
| Googlebot | answer index | 0.0% | 0 | 63 | |
| Bingbot | Microsoft | answer index | 0.0% | 0 | 63 |
Block rate is the share of this segment's domains holding a policy record for that agent whose robots.txt disallows it at the site root, shown with its numerator and denominator. Denominators differ between agents because a record only exists once a domain has been measured for it. A domain serving no robots.txt counts as allowing everything, which is what the standard specifies.
Most accessible
| Domain | Score | Engines blocked | llms.txt | |
|---|---|---|---|---|
| 1 | kiro.dev | 96 | 0 | yes |
| 2 | aikido.dev | 96 | 0 | yes |
| 3 | warp.dev | 96 | 0 | yes |
| 4 | k8slens.dev | 96 | 0 | yes |
| 5 | selenium.dev | 96 | 0 | yes |
| 6 | daily.dev | 94 | 0 | yes |
| 7 | composio.dev | 93 | 0 | yes |
| 8 | ngrok-free.dev | 91 | 0 | no |
| 9 | fga.dev | 90 | 0 | no |
| 10 | hashnode.dev | 88 | 0 | no |
| 11 | herdr.dev | 88 | 0 | yes |
| 12 | expo.dev | 88 | 0 | no |
| 13 | pub.dev | 87 | 0 | no |
| 14 | reactnative.dev | 85 | 0 | yes |
| 15 | web.dev | 83 | 0 | no |
Highest AI access scores among the 63 domains measured in this segment.
Least accessible
| Domain | Score | Engines blocked | llms.txt | |
|---|---|---|---|---|
| 1 | gunbark.dev | 36 | 18 | no |
| 2 | globed.dev | 45 | 0 | no |
| 3 | tevas.dev | 45 | 0 | no |
| 4 | deeplink.dev | 46 | 0 | no |
| 5 | github.dev | 46 | 0 | no |
| 6 | bls.dev | 48 | 0 | no |
| 7 | errortrace.dev | 49 | 0 | no |
| 8 | enot.dev | 50 | 0 | no |
| 9 | mgaru.dev | 51 | 0 | yes |
| 10 | deps.dev | 52 | 0 | no |
| 11 | floors.dev | 54 | 0 | no |
| 12 | namesvr.dev | 56 | 0 | no |
| 13 | pixeldrain.dev | 56 | 0 | no |
| 14 | zenn.dev | 57 | 0 | no |
| 15 | webtorrent.dev | 57 | 0 | no |
Lowest AI access scores among the same 63 domains.
Compare with the rest of the grouping top-level domain segments
The largest 12 other segments in this grouping, ordered by size. The share beside each is that segment's own answer-engine block rate over its own count. Full table on the segments index.
What this segment does not tell you read this before citing it
A top-level domain is a weak proxy for jurisdiction. Most of them are open to any registrant anywhere, so a domain under a country-code registry may be operated from outside that country by a company no local law reaches. Read these rows as differences between registration communities and the publishing cultures inside them, not as differences between legal regimes.
Measure your own domain
The audit that produced every number on this page runs on any domain in about six seconds: robots.txt resolved for 48 agents, then live requests sent as four of them to see whether the edge agrees with the file.