Live measurement, not robots.txt

Which edge refuses AI crawlers for you

A permissive robots.txt means nothing if the layer in front of your origin returns 403 to anything that looks like a bot. This page measures what actually happens: a real request sent as GPTBot, OAI-SearchBot, PerplexityBot and ClaudeBot from a datacentre address, grouped by the edge provider identified from response headers.

Of the 7,840 sites where an edge provider could be identified, sites behind Sucuri refuse an AI user agent most often on repeat probes, 36.5 percent of 52 measured, against 19.2 percent across the corpus as a whole.

The measurement

Of 7,840 domains where the edge provider could be identified from response headers, 36.5% of the 52 sitting behind Sucuri refuse a request carrying an AI crawler user agent, against 19.2% across the corpus. Each refusal was reproduced on two consecutive independent scans from different locations; single-observation refusals are excluded because roughly one in eleven does not reproduce.

Measured 2026-10-08 by direct request to every domain in the corpus. Method: how this is measured. Reuse under CC BY 4.0 with attribution.

Cite as: Crawl Census, "Which edge refuses AI crawlers", measured 2026-10-08. https://crawlcensus.com/edge

7,840
sites with an identified edge
1,063
confirmed refusals
14%
of identified sites
22,346
edge not identifiable
What "confirmed" means here. A single refusal is weak evidence: datacentre IP reputation, a momentary rate limit or a geographic rule can all produce one. A site is only counted as confirmed when it refused an AI user agent on two or more consecutive independent scans, run at different times from different edge locations. The single-scan column is shown beside it so the gap between the two is visible rather than hidden.
The false-positive rate, measured. Re-probing every site that refused on a first pass, roughly one in eleven did not refuse again. Those sites are excluded from the confirmed column. That is the error rate any single-probe study of AI crawler blocking carries silently, including the first pass of this one.

Refusal rate by edge provider confirmed on repeat probes

EdgeSitesConfirmed refusalRefused at least onceBlock an answer engine in robots.txtMean score
Sucuri 52 36.5%
19 of 52
40.4%
21 of 52
2% 75.8
LiteSpeed 339 36.0%
122 of 339
38.3%
130 of 339
2% 81.9
Akamai 814 24.9%
203 of 814
26.8%
218 of 814
11% 72.9
Kinsta 121 19.8%
24 of 121
20.7%
25 of 121
0% 87.7
Fastly 1,431 12.2%
175 of 1,431
14.3%
205 of 1,431
13% 77.5
CloudFront 3,544 11.8%
419 of 3,544
14.8%
524 of 3,544
11% 74.2
Vercel 686 8.2%
56 of 686
9.2%
63 of 686
3% 81.6
Azure Front Door 291 6.9%
20 of 291
7.2%
21 of 291
8% 75.5
Netlify 226 5.3%
12 of 226
5.8%
13 of 226
2% 80.4
Imperva 289 4.2%
12 of 289
5.9%
17 of 289
3% 67.9
Google Cloud 47 2.1%
1 of 47
2.1%
1 of 47
2% 63.7
bar chartSucuri (n=52)36.5%LiteSpeed (n=339)36%Akamai (n=814)24.9%Kinsta (n=121)19.8%Fastly (n=1431)12.2%CloudFront (n=3544)11.8%Vercel (n=686)8.2%Azure Front Door (n=291)6.9%Netlify (n=226)5.3%Imperva (n=289)4.2%Google Cloud (n=47)2.1%

How the edge is identified and why most cannot be

Only headers that survive the trip are trusted: x-vercel-id, x-amz-cf-id, fastly-restarts, x-akamai-request-id, x-nf-request-id, x-sucuri-id, x-iinfo, x-litespeed-cache and their siblings. The server header and cf-ray are deliberately ignored, because a request made from a Cloudflare Worker comes back carrying both regardless of what the origin actually runs. Trusting them would have reported every site on the web as Cloudflare-fronted, which is exactly the error this page exists to avoid making.

The consequence is honest but limited coverage: 22,346 measured sites expose no identifying header and are excluded entirely rather than guessed at. No provider is named unless its own header named it.

Full method on the methodology page. Underlying rows in data.json, field edge.

If your provider is on this list what to do

Being behind a provider with a high rate does not mean you are blocking anything. It usually means a managed bot rule set is enabled by default and nobody chose it deliberately. Check your own domain, then allow the retrieval agents you want citing you while leaving the rest as they are.

The report names every agent that was refused and the status it received. The guide covers detection and the fix per provider.