Blocklist · Timpi

Sites that block Timpibot

816 of the 30,338 domains we have measured, 3%, disallow timpibot at the site root in robots.txt. Timpibot collects content that may be used to train a model.

Every domain here has withheld its content from Timpi's training crawl. That is a separate decision from appearing in the operator's answers, which is governed by a different token.

The measurement

816 of 30,338 measured domains (3%) disallow timpibot at the site root in robots.txt. Every domain named on this page was fetched directly and its robots.txt parsed against RFC 9309; none of it is inferred from a third-party index.

Measured 2026-08-25 by direct request to every domain in the corpus. Method: how this is measured. Reuse under CC BY 4.0 with attribution.

Cite as: Crawl Census, "Sites that block Timpibot", measured 2026-08-25. https://crawlcensus.com/blocked/timpibot

The list 816 domains

RankDomainTLDAI access score
49,986 echo-news.co.uk .uk 71
50,054 tekstowo.pl .pl 74
50,185 miricanvas.com .com 89
50,267 worcesternews.co.uk .uk 71
50,278 burlingtonfreepress.com .com 74
50,338 ohtuleht.ee .ee 64
50,342 thenewslens.com .com 70
50,367 blipfoto.com .com 57
50,520 womansworld.com .com 81
50,598 dailycal.org .org 74
50,684 gq.com.mx .mx 64
50,916 motorcycle.com .com 83
50,919 japantravel.com .com 79
50,933 nwaonline.com .com 57
50,982 nrtool.st .st 46
51,009 treasure-maps.com .com 51
PreviousPage 5 of 5

Rank is the domain's position in the popularity list used to seed the corpus; a dash means it was scanned on request rather than seeded. Score is the site's overall AI access score, which blocking depresses but does not by itself determine.

Check a specific domain

See also the Timpibot reference, the full census, and the JSON endpoint.