Independent measurement · 30,185 sites on record

Find out what AI can actually read
on your website.

Answer engines do not see the page your visitors see. They see whatever survives robots.txt, your edge firewall, and a fetch with no JavaScript. Crawl Census measures exactly that, for any domain, in about six seconds.

Free, no account. Reports are public so they can be cited and linked.

30,185
domains measured
7%
block at least one answer engine
12%
block at least one training crawler
14%
publish an llms.txt

What gets blocked, crawler by crawler share of measured sites

bar chartGPTBot7%CCBot7%Bytespider6%ClaudeBot6%Meta-ExternalAgent5%Amazonbot5%Google-Extended5%Applebot-Extended5%anthropic-ai4%omgilibot4%Diffbot4%cohere-ai4%omgili4%SemrushBot-OCOB4%

Percentage of measured sites whose robots.txt disallows each crawler at the site root. Click a bar for the crawler's page and the full blocklist.

Recent policy changes

Full change log · RSS

Recently scanned

DomainScoreAnswer enginesScanned
education.gov.in 30 open 1 minute ago
funeral-notices.co.uk 67 open 1 minute ago
givc.ru 67 open 1 minute ago
shopsy.in 88 open 1 minute ago
lindependant.fr 86 3 blocked 1 minute ago
carwale.com 72 open 1 minute ago
jito.wtf 70 open 1 minute ago
ridewithvia.com 86 open 1 minute ago
freeones.com 71 5 blocked 1 minute ago
tycsports.com 84 open 1 minute ago

Full rankings

What the score measures 100 points

Reach · 40

Can the crawler get the bytes at all? robots.txt groups for 39 tracked agents, wildcard rules, crawl-delay, meta and header directives, plus live requests sent as GPTBot, OAI-SearchBot, PerplexityBot and ClaudeBot to catch firewall blocking that robots.txt never mentions.

Readability · 25

Is there text in the HTML? Word count and text-to-markup ratio measured with JavaScript disabled, main and article landmarks, heading structure, title and description quality.

Structure · 20

Can a machine tell what the page is about? JSON-LD validity and entity coverage, canonical URL, sitemap, and the tables, lists and question headings that answer engines lift verbatim.

Attribution · 15

Can it credit you? Author and organization entities, sameAs links, dateModified freshness, declared licence, and a conforming llms.txt.

Every check is documented on the methodology page, including the exact thresholds and how each one is measured.

Put it in your pipeline free

GitHub Action

Fail the build when a deploy makes your site less readable to AI. Runs the same audit, writes a job summary, annotates the failing checks.

- uses: taylorsmithgg/ai-access-check@v1
  with:
    domain: example.com
    min-score: 70
Repository
MCP server

Let an agent scan a site and query the census directly over streamable HTTP.

claude mcp add --transport http   crawl-census https://crawlcensus.com/mcp
Tools and schemas
JSON API

Every measurement this site publishes is available as JSON, CORS open, CC BY 4.0.

curl https://crawlcensus.com/api/v1/site/example.com
Endpoints

Why this is worth measuring context

Blocking a training crawler and blocking an answer engine are different decisions with different consequences, and most sites make them by accident with a single copied robots.txt block. GPTBot collects text that may train a model. OAI-SearchBot builds the index that decides whether ChatGPT cites you. Disallow both and you have not protected anything, you have removed yourself from the results while your competitor stays in them.

The second failure is quieter. A permissive robots.txt means nothing if the edge returns 403 to any user agent containing "bot", which is the default behaviour of several managed rule sets. That is why every report here includes live requests sent under real crawler user agents, not just a reading of the file.

The third is structural. Most AI fetchers do not execute JavaScript. A single-page app that renders its content client-side ships an empty shell to the crawler no matter how permissive the rules are, and the engine summarises the shell.

Crawl Census re-measures continuously and keeps the history, so the record shows not just who is open today but who changed their mind, and when.