Live measurement · a price, not a refusal

Sites that charge AI crawlers

Most coverage of AI crawling treats the question as binary: a site either allows a crawler or blocks it. A third answer is now measurable in the wild. These domains respond to an AI crawler user agent with HTTP 402 Payment Required - the status reserved for exactly this, and the one Cloudflare's pay-per-crawl returns. They are not refusing to be read. They are quoting a rate.

The measurement

252 of 29,277 measured domains answer an AI crawler user agent with HTTP 402 Payment Required rather than serving the page or refusing it. For comparison, 5,462 refuse outright. Counting 402 as a block, as a status-code-only method would, overstates refusal by 4% of the total.

Measured 2026-08-24 by direct request to every domain in the corpus. Method: how this is measured. Reuse under CC BY 4.0 with attribution.

Cite as: Crawl Census, "Sites that charge AI crawlers", measured 2026-08-24. https://crawlcensus.com/charging

252
domains quoting a price
161
charge every agent alike
47
block training, charge retrieval
4%
of non-serving responses are a price

The split policy 47 domains

The most interesting response is not a single status code but a pair of them. 47 domains refuse the training crawlers outright and quote a price to the retrieval crawlers, in the same second, from the same address.

A request as GPTBot or ClaudeBot returns 403. A request as OAI-SearchBot or PerplexityBot returns 402. The operator is stating a commercial position with precision: do not use this to train a model, and you may cite it in an answer if you pay. That distinction is the substance of every publisher negotiation currently running, and here it is implemented in production and measurable from outside.

No robots.txt file can express this. A policy file can allow or disallow a token; it cannot quote a rate, and it cannot make the answer depend on what the crawler intends to do with the page. This is why a robots.txt-only study cannot see the split at all.

Every domain measured returning 402 the list

DomainRankAI access scoreEdge
usatoday.com#55275not identified
independent.co.uk#66988Fastly
apnews.com#78784not identified
theatlantic.com#98079Fastly
clarin.com#1,13682not identified
pimpbunny.com#1,17761not identified
toyhou.se#1,43954not identified
hdporncomics.com#1,45085not identified
ijavhd.com#1,59471not identified
about.com#1,94366not identified
mirror.co.uk#2,11547CloudFront
variety.com#2,11880not identified
slate.com#2,34680Fastly
rollingstone.com#2,35283not identified
lww.com#2,36763not identified
hollywoodreporter.com#2,58179not identified
billboard.com#3,25378not identified
deadline.com#3,41881not identified
politico.eu#3,59562not identified
stern.de#4,08383Akamai
japantimes.co.jp#4,22783CloudFront
snopes.com#4,33085not identified
blender.org#4,33465not identified
gamer.com.tw#5,86674not identified
azcentral.com#6,08969not identified
wwd.com#6,30878not identified
bbcgoodfood.com#6,35384Fastly
freep.com#6,35673not identified
telegram.com#6,51273not identified
gbnews.com#6,60088Fastly
lavoz.com.ar#7,08775not identified
superuser.com#7,10382not identified
sport.de#7,18284CloudFront
manchestereveningnews.co.uk#7,25047CloudFront
nzz.ch#7,26876not identified
rae.es#7,34963not identified
ole.com.ar#7,38381not identified
rd.com#7,44991not identified
giallozafferano.it#7,46887not identified
indiewire.com#7,51075not identified
calciomercato.com#7,93781not identified
tudocelular.com#8,05788Fastly
my-personaltrainer.it#8,12779not identified
walesonline.co.uk#8,29247CloudFront
voetbalzone.nl#8,66481not identified
bangkokpost.com#8,90789not identified
pomorska.pl#8,94477not identified
sorrisi.com#9,24181not identified
detroitnews.com#9,33968not identified
mintmanga.one#9,38492not identified
jsonline.com#9,59973not identified
reviewjournal.com#10,49182not identified
dailystar.co.uk#10,62147CloudFront
liverpoolecho.co.uk#10,64647CloudFront
soap2day.day#11,02259not identified
guidestar.org#11,28659not identified
ballotpedia.org#11,29546CloudFront
houstonchronicle.com#11,46554Fastly
zoznam.sk#11,61774not identified
uottawa.ca#11,64178not identified
dailyrecord.co.uk#11,75147CloudFront
ovid.com#11,97267not identified
acog.org#12,26272not identified
tasteofhome.com#12,28491not identified
indystar.com#12,53373not identified
irishexaminer.com#12,86489not identified
radiopaedia.org#12,92754not identified
nolo.com#13,04589not identified
oas.org#13,08955not identified
skyscrapercity.com#13,45143Fastly
the-independent.com#13,60790Fastly
dispatch.com#13,66769not identified
sheknows.com#13,79282not identified
tennessean.com#13,87974not identified
macrotrends.net#14,02151not identified
observador.pt#14,05577not identified
wccftech.com#14,07976not identified
postcodebase.com#14,41079not identified
birminghammail.co.uk#15,26047CloudFront
desmoinesregister.com#15,40174not identified
cincinnati.com#15,55774not identified
rebrickable.com#15,62987not identified
fibaro.com#16,05366not identified
mobiauto.com.br#16,54292not identified
bwfbadminton.com#16,93556not identified
familyhandyman.com#16,98190not identified
northjersey.com#17,04274not identified
montrealgazette.com#17,09379not identified
heavyfetish.com#17,18962not identified
harpers.org#17,22477not identified
palmbeachpost.com#17,46973not identified
manilatimes.net#17,48079CloudFront
xfreehd.com#17,52075not identified
hdblog.it#17,53786Fastly
autocar.co.uk#17,73684CloudFront
aktualne.cz#17,82790not identified
oklahoman.com#18,20874not identified
lasvegassun.com#18,53978not identified
cas.sk#18,65284not identified
courier-journal.com#18,88374not identified
appstate.edu#19,71374not identified
echodnia.eu#19,73879not identified
gazetawroclawska.pl#19,87079not identified
geo.de#19,88282Akamai
chroniclelive.co.uk#20,08047CloudFront
vhlcentral.com#20,14459not identified
sacred-texts.com#20,48476CloudFront
artnews.com#20,57545not identified
nto.pl#20,76479not identified
bps.org.uk#21,04469not identified
fnal.gov#21,05072not identified
telemagazyn.pl#21,16181not identified
save-free.com#21,27891not identified
zmescience.com#21,79692not identified
airliners.net#22,21644Fastly
fortebet.ug#22,67743not identified
shanghaifantasy.com#22,67858not identified
app.com#23,08469not identified
rhizome.org#23,41559not identified
dailypost.ng#23,84687not identified
knoxnews.com#24,10269not identified
jacksonville.com#24,19074not identified
irishmirror.ie#24,41347CloudFront
humanite.fr#24,73387not identified
bristolpost.co.uk#24,80247CloudFront
myrecipes.com#25,13377not identified
consequence.net#25,67882not identified
lexicanum.com#25,80169not identified
alstom.com#25,82465not identified
delawareonline.com#25,83974not identified
thoughtcatalog.com#25,88582not identified
heraldtribune.com#25,90373not identified
people.inc#26,00564not identified
allthatsinteresting.com#26,03358not identified
democratandchronicle.com#26,11473not identified
spox.com#26,56782not identified
tallahassee.com#26,85374not identified
fd.nl#26,93573not identified
themirror.com#27,24347CloudFront
amny.com#27,25679not identified
stylecaster.com#27,30664not identified
americanbanker.com#27,45680CloudFront
texasattorneygeneral.gov#27,73763not identified
asiasociety.org#27,79672Fastly
centrum.sk#27,79773not identified
blogg.se#28,69781not identified
floridatoday.com#29,14274not identified
yamareco.com#29,14965not identified
kompas.id#29,39069not identified
providencejournal.com#29,82173not identified
petitfute.com#30,23368not identified
clarionledger.com#30,23774not identified
wickedlocal.com#30,70375not identified
artforum.com#30,71274not identified
commercialappeal.com#30,98674not identified
coventrytelegraph.net#31,38147CloudFront
lohud.com#31,54974not identified
avsforum.com#31,85043Fastly
heraldnet.com#31,89486not identified
fapcat.com#31,93664not identified
devonlive.com#32,15147CloudFront
teamblind.com#32,17472CloudFront
gloswielkopolski.pl#32,37379not identified
rgj.com#32,38873not identified
chittorgarh.com#32,73660not identified
dailypost.co.uk#33,06347CloudFront
desertsun.com#33,33274not identified
nottinghampost.com#33,56947CloudFront
columbian.com#34,37575not identified
mylondon.news#34,44947CloudFront
stylist.co.uk#34,60363CloudFront
vibe.com#34,73780not identified
ooyyo.com#34,84668not identified
lensa.com#34,94683not identified
naplesnews.com#34,97373not identified
historyextra.com#35,19882Fastly
everyeye.it#35,29387not identified
totvs.com#35,48688not identified
news-journalonline.com#35,54373not identified
mediaindonesia.com#35,55780not identified
dotdashmeredith.com#35,82864not identified
football.london#35,87747CloudFront
icheckmovies.com#35,95056not identified
redflagdeals.com#35,95344Fastly
chalkbeat.org#36,31874Akamai
asahq.org#37,03266not identified
bloodyelbow.com#37,31788not identified
pluska.sk#37,54984not identified
youbianku.com#37,60780not identified
leicestermercury.co.uk#37,67447CloudFront
capital.de#37,72283Akamai
news-press.com#38,14873not identified
goldderby.com#38,38874not identified
rollitup.org#38,83083not identified
shawlocal.com#39,12671Akamai
thecentersquare.com#39,21078not identified
vwvortex.com#39,47143Fastly
hulldailymail.co.uk#39,58747CloudFront
belfastlive.co.uk#40,01347CloudFront
saudigazette.com.sa#40,06586CloudFront
newsok.com#40,10982not identified
detnews.com#40,26275not identified
emedicinehealth.com#40,34777not identified
elpasotimes.com#40,79673not identified
stokesentinel.co.uk#41,10447CloudFront
cjonline.com#41,64773not identified
sarenza.com#41,72776CloudFront
thehealthy.com#41,86589not identified
peopleenespanol.com#41,91877not identified
theledger.com#41,95874not identified
tcpalm.com#42,16073not identified
rcgroups.com#42,26041Fastly
advrider.com#42,26243Fastly
cars.co.za#42,65270not identified
wishtv.com#42,72881Fastly
polaris.me#42,72950not identified
nosofiles.com#42,95738not identified
findmeglutenfree.com#42,97963not identified
derbytelegraph.co.uk#43,43747CloudFront
filefactory.com#43,54662not identified
pbtech.co.nz#43,67182not identified
investmentnews.com#43,92086not identified
nctm.org#43,96166not identified
chargemap.com#43,96391Fastly
whatcar.com#43,99492CloudFront
cambridge-news.co.uk#44,04147CloudFront
giantfreakinrobot.com#44,31888not identified
totvs.com.br#44,36085not identified
cornwalllive.com#44,94647CloudFront
gulbenkian.pt#45,24777not identified
traderie.com#45,28347not identified
lihkg.com#45,38650not identified
gloucestershirelive.co.uk#45,85047CloudFront
nyheter24.se#46,15187not identified
gazettelive.co.uk#46,16447CloudFront
gardenersworld.com#46,18183Fastly
registerguard.com#46,23166Fastly
consequenceofsound.net#46,33982not identified
pjstar.com#46,34266Fastly
pnj.com#46,36062Fastly
plymouthherald.co.uk#46,76547CloudFront
examinerlive.co.uk#46,91847CloudFront
teslamotorsclub.com#46,97272not identified
fixya.com#46,99043Fastly
meredith.com#47,20866not identified
vcstar.com#47,88462Fastly
ukclimbing.com#48,00470not identified
wsls.com#48,38086Akamai
lesoleil.com#48,40465Akamai
argusleader.com#48,55062Fastly
sj-r.com#49,50164Fastly
montgomeryadvertiser.com#49,57462Fastly

A price nobody states 242 of 252 say nothing

A 402 is only useful to a crawler if it says what it wants. Every response from these domains was captured and searched for the headers a pay-per-crawl scheme would use to state terms - crawler-price, www-authenticate, link, signature-agent and others.

242 of 252 domains answering HTTP 402 state nothing a machine can act on. The 10 that send anything at all send retry-after: 0, which tells a crawler to try again immediately and says nothing about money.

So the status code is being used as a signal rather than as a transaction. An operator publishing 402 is saying "not for free" to an audience that has no way to pay, and the crawler on the other side has no machine-readable path from the refusal to the content. That gap is worth naming precisely, because it is the difference between a working market and a locked door with a price tag nobody can read. Whatever terms exist are arranged out of band, between companies, which is exactly what the status code was meant to replace.

Crawl Census reports this rather than guessing at it: the preflight endpoint returns any captured terms verbatim on a pay verdict, and says plainly when there are none.

Why this is a separate category method

A 403 says the operator does not want this crawler to read the page. A 402 says the operator will serve it, on terms. Those are opposite commercial positions and they produce identical rows in any study that buckets by "non-200 response", which is how most crawler-blocking figures are produced.

This distinction is only visible to a measurement that sends a real request. It cannot be derived from robots.txt, because a site quoting a price typically keeps a permissive robots.txt: the transaction happens at the edge, not in a policy file. Of the domains listed here, the majority publish robots.txt rules that allow the very agents they charge.

Status 402 was reserved for this purpose in the original HTTP specification and went essentially unused for thirty years. Its appearance in production on prominent publishers is a measurable change in how the web is monetised, not a configuration error.