Sites that block Googlebot
16 of the 27,307 domains we have measured, 0.1%, disallow googlebot at the site root in robots.txt. Googlebot builds the index that decides whether the operator's assistant cites a page.
Every domain here has removed itself from Google's answers. Blocking a retrieval agent withholds nothing from model training, because training uses a different token.
16 of 27,307 measured domains (0.1%) disallow googlebot at the site root in robots.txt. Every domain named on this page was fetched directly and its robots.txt parsed against RFC 9309; none of it is inferred from a third-party index.
Measured 2026-08-23 by direct request to every domain in the corpus. Method: how this is measured. Reuse under CC BY 4.0 with attribution.
Cite as: Crawl Census, "Sites that block Googlebot", measured 2026-08-23. https://crawlcensus.com/blocked/googlebot
The list 16 domains
| Rank | Domain | TLD | AI access score |
|---|---|---|---|
| 1,754 | ad-contents.jp | .jp | 52 |
| 4,200 | disdarrummers.shop | .shop | 77 |
| 4,619 | dextrodedenda.top | .top | 77 |
| 5,597 | feinidi.com | .com | 80 |
| 7,228 | ishe88.com | .com | 80 |
| 10,172 | dyz9.top | .top | 80 |
| 16,122 | page-stats.de | .de | 49 |
| 20,137 | 3bb.co.th | .th | 71 |
| 23,458 | inveragil.com | .com | 51 |
| 25,270 | startersites.io | .io | 46 |
| 25,826 | cnnic.cn | .cn | 59 |
| 25,857 | ly.com | .com | 68 |
| 29,683 | officeplus.cn | .cn | 60 |
| 30,868 | sosalkino.best | .best | 69 |
| 39,484 | brighteon.com | .com | 55 |
| 44,129 | eeo.cn | .cn | 54 |
Rank is the domain's position in the popularity list used to seed the corpus; a dash means it was scanned on request rather than seeded. Score is the site's overall AI access score, which blocking depresses but does not by itself determine.
Check a specific domain
See also the Googlebot reference, the full census, and the JSON endpoint.