Sites that block OAI-SearchBot
649 of the 29,415 domains we have measured, 2%, disallow oai-searchbot at the site root in robots.txt. OAI-SearchBot builds the index that decides whether the operator's assistant cites a page.
Every domain here has removed itself from OpenAI's answers. Blocking a retrieval agent withholds nothing from model training, because training uses a different token.
649 of 29,415 measured domains (2%) disallow oai-searchbot at the site root in robots.txt. Every domain named on this page was fetched directly and its robots.txt parsed against RFC 9309; none of it is inferred from a third-party index.
Measured 2026-08-24 by direct request to every domain in the corpus. Method: how this is measured. Reuse under CC BY 4.0 with attribution.
Cite as: Crawl Census, "Sites that block OAI-SearchBot", measured 2026-08-24. https://crawlcensus.com/blocked/oai-searchbot
The list 649 domains
Rank is the domain's position in the popularity list used to seed the corpus; a dash means it was scanned on request rather than seeded. Score is the site's overall AI access score, which blocking depresses but does not by itself determine.
Check a specific domain
See also the OAI-SearchBot reference, the full census, and the JSON endpoint.