Sites that block Meta-ExternalAgent
2,699 of the 30,351 domains we have measured, 9%, disallow meta-externalagent at the site root in robots.txt. Meta-ExternalAgent collects content that may be used to train a model.
Every domain here has withheld its content from Meta's training crawl. That is a separate decision from appearing in the operator's answers, which is governed by a different token.
2,699 of 30,351 measured domains (9%) disallow meta-externalagent at the site root in robots.txt. Every domain named on this page was fetched directly and its robots.txt parsed against RFC 9309; none of it is inferred from a third-party index.
Measured 2026-08-25 by direct request to every domain in the corpus. Method: how this is measured. Reuse under CC BY 4.0 with attribution.
Cite as: Crawl Census, "Sites that block Meta-ExternalAgent", measured 2026-08-25. https://crawlcensus.com/blocked/meta-externalagent
The list 2,699 domains
Rank is the domain's position in the popularity list used to seed the corpus; a dash means it was scanned on request rather than seeded. Score is the site's overall AI access score, which blocking depresses but does not by itself determine.
Check a specific domain
See also the Meta-ExternalAgent reference, the full census, and the JSON endpoint.