AI crawler · registry

Applebot

Apple · Answer index · robots.txt token applebot

honors robots.txtMeasured across 30,177 sites with a robots.txt policy record for this agent, out of 30,185 on record.
The measurement

1% of 30,177 measured domains disallow applebot in robots.txt at the site root. Separately, and independently of robots.txt, 17.4% of 30,185 reachable domains refuse an AI crawler user agent at the network edge on two consecutive probes from different locations - a refusal the site's own robots.txt does not declare and its owner often has not chosen.

Measured 2026-10-06 by direct request to every domain in the corpus. Method: how this is measured. Reuse under CC BY 4.0 with attribution.

Cite as: Crawl Census, "Applebot blocking rate", measured 2026-10-06. https://crawlcensus.com/bot/applebot

1%
of measured sites block it
343
sites blocking
29,834
sites allowing
yes
honors robots.txt

What it is definition

Crawls for Siri, Spotlight and Safari search; falls back to Googlebot rules and ignores Crawl-delay.

Applebot is operated by Apple and classified here as answer index. These agents build the index an answer engine queries at question time. Disallowing one removes the site from those answers and from the citations that link back to it. Training is governed by separate tokens, so this is the expensive block, not the protective one.

Disallowing applebot removes the site from the index Siri, Spotlight and Safari search queries when it answers a question. The engine can then neither quote nor link the page, which is a traffic decision rather than a licensing one. Model training is governed by separate tokens.

robots.txt token
applebot
Operator
Apple
Purpose
Answer index
Honors robots.txt
yes
Documentation
https://support.apple.com/en-us/119829

Identity on the wire exact string

User-agent string
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)
Verify your own edge is not refusing it
curl -A 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)' https://example.com/

Replace the host with your own. A 403, a challenge page or an empty body means the edge is blocking the agent whatever robots.txt says.

How to allow it, how to block it robots.txt

Allow it
User-agent: applebot
Allow: /

Absent any rule, the agent is already allowed, so this group only matters when your * group disallows something. Naming the agent replaces the * group for it entirely, so repeat here every disallow you still want to apply to it.

Block it
User-agent: applebot
Disallow: /

Disallowing applebot removes the site from the index Siri, Spotlight and Safari search queries when it answers a question. The engine can then neither quote nor link the page, which is a traffic decision rather than a licensing one. Model training is governed by separate tokens.

Sites that block it 60 shown

DomainAI access scoreRank
msn.com 41 62
usatoday.com 75 552
uol.com.br 82 571
cnet.com 79 662
repubblica.it 85 942
corriere.it 76 955
theconversation.com 62 1,019
franceinfo.fr 62 1,098
idnes.cz 80 1,104
furaffinity.net 39 1,223
dw.com 42 1,305
eenadu.net 67 1,456
zdnet.com 72 1,831
themoviedb.org 71 1,984
ign.com 67 1,993
mashable.com 61 2,010
mparticle.com 64 2,258
tmdb.org 71 2,656
chicagotribune.com 81 2,786
imageshack.us 31 2,975
lifehacker.com 65 3,342
stuff.co.nz 54 3,588
politico.eu 62 3,595
nydailynews.com 79 3,621
ilsole24ore.com 77 4,070
kompas.com 77 4,161
zoominfo.com 87 4,190
tvtropes.org 44 4,304
snopes.com 85 4,330
lesechos.fr 49 4,378
biorxiv.org 44 4,532
lastampa.it 82 4,592
tomsguide.com 82 4,787
dr.dk 81 4,987
imageshack.com 31 5,037
everydayhealth.com 60 5,146
mercurynews.com 79 5,158
lapresse.ca 65 5,205
chinatimes.com 80 5,297
nimo.tv 66 5,711
francetvinfo.fr 62 5,714
tv2.no 70 5,775
runescape.wiki 81 5,846
bustle.com 85 5,909
ilmessaggero.it 62 6,004
delfi.lt 80 6,064
futura-sciences.com 93 6,088
azcentral.com 74 6,089
denverpost.com 79 6,125
eurogamer.net 57 6,246
freep.com 73 6,356
parismatch.com 74 6,396
telegram.com 73 6,512
sandiegouniontribune.com 78 6,513
maxroll.gg 59 6,816
nicematin.com 77 6,948
military.com 85 7,109
rp-online.de 79 7,119
baltimoresun.com 73 7,275
bhaskar.com 75 7,363

See every measured domain that blocks Applebot

This list is served from a cache rebuilt just now. It is discarded the moment a rescan records a policy change for this agent, so a site that has just changed its robots.txt will drop off it within one scan rather than waiting for the cache to expire.

Compare Apple

Apple runs 2 agents in this registry, and they do different jobs. Blocking one says nothing about the others.

Common questions answered

Does blocking Applebot remove me from Siri, Spotlight and Safari search?

Yes, in effect. Applebot builds the index Siri, Spotlight and Safari search reads at question time, so disallowing it takes the site out of those answers and out of the citations that link back to it. This is the block that costs traffic.

How do I verify Applebot requests are genuine?

Run a reverse DNS lookup on the source address. It must resolve to a host under applebot.apple.com, and a forward lookup of that host must return the same address.

Check your own site free

See whether your robots.txt admits Applebot today, and whether your edge agrees with it.