AI crawler · registry

Meta-WebIndexer

Meta · Answer index · robots.txt token meta-webindexer

honors robots.txtMeasured across 3,958 sites with a robots.txt policy record for this agent, out of 7,068 on record.
2%
of measured sites block it
81
sites blocking
3,877
sites allowing
yes
honors robots.txt

What it is definition

Indexes pages so Meta AI can cite and link them in its search answers.

Meta-WebIndexer is operated by Meta and classified here as answer index. These agents build the index an answer engine queries at question time. Disallowing one removes the site from those answers and from the citations that link back to it. Training is governed by separate tokens, so this is the expensive block, not the protective one.

Disallowing meta-webindexer removes the site from the index Meta's AI products queries when it answers a question. The engine can then neither quote nor link the page, which is a traffic decision rather than a licensing one. Model training is governed by separate tokens.

robots.txt token
meta-webindexer
Operator
Meta
Purpose
Answer index
Honors robots.txt
yes
Documentation
https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/

Identity on the wire exact string

User-agent string
meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)
Verify your own edge is not refusing it
curl -A 'meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)' https://example.com/

Replace the host with your own. A 403, a challenge page or an empty body means the edge is blocking the agent whatever robots.txt says.

How to allow it, how to block it robots.txt

Allow it
User-agent: meta-webindexer
Allow: /

Absent any rule, the agent is already allowed, so this group only matters when your * group disallows something. Naming the agent replaces the * group for it entirely, so repeat here every disallow you still want to apply to it.

Block it
User-agent: meta-webindexer
Disallow: /

Disallowing meta-webindexer removes the site from the index Meta's AI products queries when it answers a question. The engine can then neither quote nor link the page, which is a traffic decision rather than a licensing one. Model training is governed by separate tokens.

Sites that block it 60 shown

DomainAI access scoreRank
amazon.com 33 26
amazonvideo.com 55 121
nytimes.com 67 157
amazon.co.uk 33 318
amazon.de 33 377
amazon.co.jp 33 494
wired.com 74 533
usatoday.com 55 552
amazon.fr 33 609
amazon.co.za 33 623
amazon.ca 33 635
amazon.in 34 659
cnet.com 75 662
lemonde.fr 72 671
primevideo.com 55 682
amazon.it 33 706
amazon.es 33 725
people.com 38 801
amazon.com.au 33 812
amazon.com.br 33 833
investopedia.com 38 913
theconversation.com 52 1,019
amazon.com.mx 33 1,032
bfmtv.com 71 1,073
furaffinity.net 41 1,223
iltalehti.fi 50 1,409
eenadu.net 73 1,456
newyorker.com 71 1,793
zdnet.com 61 1,831
wikihow.com 72 1,835
arstechnica.com 74 1,847
ign.com 64 1,993
amazon.eg 34 2,006
mashable.com 62 2,010
techradar.com 83 2,215
congress.gov 33 2,293
pcmag.com 65 2,423
livescience.com 81 2,585
theglobeandmail.com 59 2,659
vanityfair.it 74 2,899
imageshack.us 34 2,975
vogue.com 75 2,989
allrecipes.com 38 3,184
letour.fr 69 3,210
myanimelist.net 63 3,276
tomshardware.com 83 3,331
lifehacker.com 68 3,342
amazon.se 33 3,499
amazon.ae 33 3,520
vanityfair.com 74 3,533
amazon.nl 33 3,598
space.com 79 3,873
radiofrance.fr 76 3,882
amazon.pl 33 4,097
thestar.com 65 4,107
kompas.com 76 4,161
amazon.sa 33 4,240
verywellhealth.com 38 4,285
verywellmind.com 38 4,344
southernliving.com 38 4,491

Compare Meta

Meta runs 6 agents in this registry, and they do different jobs. Blocking one says nothing about the others.

Common questions answered

Does blocking Meta-WebIndexer remove me from Meta's AI products?

Yes, in effect. Meta-WebIndexer builds the index Meta's AI products reads at question time, so disallowing it takes the site out of those answers and out of the citations that link back to it. This is the block that costs traffic.

How do I verify Meta-WebIndexer requests are genuine?

Meta's crawlers originate from AS32934, and Meta documents querying that ASN to obtain the current ranges. Map the source address to its ASN and check for AS32934. There is no published reverse DNS name.

Check your own site free

See whether your robots.txt admits Meta-WebIndexer today, and whether your edge agrees with it.