The AI crawler census
Every domain here was measured the same way: read its robots.txt, resolve the rules for 48 tracked AI agents, then send live requests as four of them and record what came back. No survey, no self-reporting. Figures update as the crawler works through the queue.
What share of the web can an AI crawler actually read the honest denominator
The headline rate elsewhere on this site, 19%, is measured across the 30,185 websites that served this crawler a page. It leaves out the 5,739 websites that refused before any question about AI could be asked - a bare 403 to an ordinary browser user agent, no robots.txt consulted, no content returned.
Excluding them understates the answer, because a site that refuses every programmatic client refuses AI crawlers most of all. Counted properly, of 35,924 websites in the corpus, 11,560 - 32% - cannot be read by an AI crawler: 5,739 refuse outright and 5,821 of the remainder refuse an AI user agent specifically.
The same exclusion could have distorted the robots.txt figures, and it does not. Those 5,739 sites refused to serve a page but still served a policy file, because robots.txt is fetched in the same pass rather than after the home page succeeds. Their stated policy is indistinguishable from everyone else's: 6.7% disallow an answer engine against 6.6% among sites that answered, and 11.4% disallow a training crawler against 11.7%. A site's written policy turns out to be independent of whether its infrastructure tolerates a programmatic client at all - which is the same split this census keeps finding between what operators decide and what their stack does by default. The robots.txt rates elsewhere on this site are therefore left on the narrower denominator, because widening it changes nothing.
Both numbers are published because they answer different questions. The narrower one is what this crawler directly measured. The wider one is what a person asking "how much of the web can a model actually read" should be told, and it is nearly twice as large. Among the sites refusing outright are openai.com, chatgpt.com and perplexity.ai, whose own products depend on crawling everyone else.
What the corpus is made of 49,557 hostnames
The seed list ranks hostnames by traffic, and a large share of the most trafficked hostnames on the internet are not websites: CDN endpoints, update services and API hosts that answer enormous numbers of DNS queries and serve no page at their apex. akamaiedge.net, cloudfront.net, windowsupdate.com and aaplimg.com all sit in the first hundred and none of them has a home page.
| Outcome | Hostnames | Share |
|---|---|---|
| Website measured | 30,185 | 60.9% |
| No website at this hostname | 12,649 | 25.5% |
| Refused before assessment | 5,739 | 11.6% |
| Other HTTP error | 984 | 2.0% |
Every rate published on this site divides by the first row only, because a hostname with no website cannot have a crawler policy. The distinction was verified rather than assumed: 32 hostnames returning a Cloudflare origin error were fetched again from a residential connection with redirects followed, and none of them served a page there either.
How much of this depends on how prominent a site is the corpus is Tranco's top 36,000
Every figure on this page describes domains drawn from the Tranco popularity ranking, positions 2 to 35,634. It is not a sample of the web, which is mostly small and obscure, and rates measured here should not be read as web-wide. The gradient below is published so the direction and size of that bias is visible rather than merely disclosed.
| Popularity band | Domains | Refuse at the edge | Block an answer engine | Publish llms.txt | Mean score |
|---|---|---|---|---|---|
| Top 1,000 | 484 | 13.8% | 12% | 17.8% | 72.2 |
| 1,000 to 10,000 | 5,191 | 20% | 9% | 14.3% | 73.3 |
| 10,000 and below | 24,404 | 19.3% | 6% | 13.4% | 73.4 |
| Unranked | 106 | 14.2% | 0.9% | 31.1% | 78.1 |
The two block types behave differently, and that is the finding. Declaring a block in robots.txt tracks prominence closely: the most visited thousand domains do it at roughly twice the rate of those ranked ten thousand and below. Refusing an AI user agent at the edge does not follow that curve at all - it peaks in the middle of the ranking, among sites large enough to sit behind a managed WAF but not large enough to have anyone tuning its bot rules. The first number is a decision. The second is mostly a default.
Across 30,185 domains measured by direct request, 7% disallow at least one answer engine in robots.txt and 12% disallow at least one training crawler. Independently of robots.txt, 19% refuse an AI crawler user agent at the network edge. 14% publish an llms.txt.
Measured 2026-10-06 by direct request to every domain in the corpus. Method: how this is measured. Reuse under CC BY 4.0 with attribution.
Cite as: Crawl Census, "The AI crawler census", measured 2026-10-06. https://crawlcensus.com/census
How the web's posture is moving share of measured sites, by day
Block rate by crawler the headline table
| Crawler | Operator | Content used for | Block rate | Sites blocking | Sites measured |
|---|---|---|---|---|---|
| GPTBot | OpenAI | training | 7.1% | 2,145 | 30,177 |
| CCBot | Common Crawl Foundation | archive | 6.8% | 2,066 | 30,177 |
| Bytespider | ByteDance | training | 6.5% | 1,950 | 30,177 |
| ClaudeBot | Anthropic | training | 5.8% | 1,743 | 30,177 |
| Meta-ExternalAgent | Meta | training | 4.7% | 1,430 | 30,177 |
| Amazonbot | Amazon | training | 4.7% | 1,411 | 30,177 |
| Google-Extended | training | 4.6% | 1,391 | 30,177 | |
| Applebot-Extended | Apple | training | 4.6% | 1,378 | 30,177 |
| anthropic-ai | Anthropic | training | 4.4% | 1,320 | 30,177 |
| omgilibot | Webz.io | training | 4.3% | 1,289 | 30,177 |
| Diffbot | Diffbot | training | 4.1% | 1,224 | 30,177 |
| cohere-ai | Cohere | live retrieval | 4.1% | 1,223 | 30,177 |
| omgili | Webz.io | training | 3.9% | 1,171 | 30,177 |
| SemrushBot-OCOB | Semrush | training | 3.7% | 1,126 | 30,177 |
| PerplexityBot | Perplexity | answer index | 3.5% | 1,054 | 30,177 |
| ChatGPT-User | OpenAI | live retrieval | 3.5% | 1,053 | 30,177 |
| FacebookBot | Meta | training | 3.2% | 979 | 30,177 |
| YouBot | You.com | answer index | 2.9% | 890 | 30,177 |
| ImagesiftBot | ImageSift (Hive) | training | 2.8% | 853 | 30,177 |
| Timpibot | Timpi | training | 2.8% | 833 | 30,177 |
| Scrapy | Zyte (open-source framework) | training | 2.5% | 745 | 30,177 |
| AI2Bot | Allen Institute for AI | training | 2.3% | 706 | 30,177 |
| OAI-SearchBot | OpenAI | answer index | 2.2% | 650 | 30,177 |
| cohere-training-data-crawler | Cohere | training | 2.1% | 620 | 30,177 |
| Meta-ExternalFetcher | Meta | live retrieval | 1.8% | 557 | 30,177 |
| DuckAssistBot | DuckDuckGo | live retrieval | 1.7% | 519 | 30,177 |
| PanguBot | Huawei | training | 1.7% | 518 | 30,177 |
| Claude-User | Anthropic | live retrieval | 1.7% | 498 | 30,177 |
| Claude-SearchBot | Anthropic | answer index | 1.6% | 484 | 30,177 |
| Perplexity-User | Perplexity | live retrieval | 1.6% | 469 | 30,177 |
| Kangaroo Bot | Kangaroo LLM | training | 1.5% | 466 | 30,177 |
| MistralAI-User | Mistral AI | live retrieval | 1.5% | 447 | 30,177 |
| Google-CloudVertexBot | live retrieval | 1.4% | 423 | 30,177 | |
| GoogleOther | training | 1.3% | 383 | 30,177 | |
| Meta-WebIndexer | Meta | answer index | 1.2% | 359 | 30,177 |
| Applebot | Apple | answer index | 1.1% | 343 | 30,177 |
| TikTokSpider | ByteDance | live retrieval | 0.8% | 255 | 30,177 |
| Amzn-SearchBot | Amazon | answer index | 0.8% | 233 | 30,177 |
| Amzn-User | Amazon | live retrieval | 0.6% | 167 | 30,177 |
| MistralAI-Training | Mistral AI | training | 0.5% | 142 | 30,177 |
| MistralAI-Index | Mistral AI | answer index | 0.3% | 81 | 30,177 |
| facebookexternalhit | Meta | live retrieval | 0.2% | 53 | 30,177 |
| Bingbot | Microsoft | answer index | 0.1% | 40 | 30,177 |
| Googlebot-News | answer index | 0.1% | 36 | 30,177 | |
| Meta-ExternalAds | Meta | training | 0.1% | 22 | 30,177 |
| msnbot | Microsoft | answer index | 0.1% | 21 | 30,177 |
| Googlebot | answer index | 0.1% | 16 | 30,177 | |
| OAI-AdsBot | OpenAI | live retrieval | 0.0% | 11 | 30,177 |
Block rate is the share of measured sites whose robots.txt disallows that agent at the site root. A site with no robots.txt counts as allowing everything, which is what the standard specifies.
Why this table's total is lower than 30,185. The headline counts every domain that answered over HTTPS. This table counts only domains that already have a stored policy record for that specific agent, which is a snapshot taken at the last daily rollup and therefore trails the live crawl. The two converge as the queue drains. Both numbers are shown rather than reconciled silently.
Grade distribution
Blocking by top-level domain
| TLD | Sites | Block an answer engine | Publish llms.txt |
|---|---|---|---|
| .com | 14,781 | 7% | 17% |
| .org | 1,666 | 6% | 9% |
| .net | 1,301 | 4% | 10% |
| .ru | 1,165 | 3% | 9% |
| .de | 825 | 16% | 7% |
| .io | 661 | 2% | 29% |
| .edu | 471 | 2% | 7% |
| .uk | 447 | 17% | 9% |
| .jp | 416 | 9% | 5% |
| .br | 374 | 6% | 13% |
| .fr | 329 | 15% | 3% |
| .in | 304 | 3% | 10% |
A TLD is a weak proxy for jurisdiction, since most are open to anyone, but the spread between them is large and consistent. Full segmentation.
Take the data free, attributed
The corpus is published under CC BY 4.0. Attribute as Source: Crawl Census (crawlcensus.com).