Every AI crawler we track
The list is split by what the operator does with the content, because that is the only division that changes the consequence of a disallow. Withholding content from model training and removing yourself from AI answers are opposite decisions, and a single copied robots.txt block usually makes both at once. Note also that a robots.txt token is not a user agent: the token is the name you write in the file, the user agent is the string that arrives on the wire, and for two of the agents below no such string exists at all.
30,185 sites on record. 48 agents tracked across 4 purposes.
Block rates, highest first top 12
Share of sites with a policy record for that agent whose robots.txt disallows it at the site root. Only agents with at least one record appear.
Model training 23 agents
These agents collect text and images that may end up in a training corpus. Disallowing one withholds new content from the next model. It does not withdraw anything collected before the disallow, and it does not remove the site from any answer engine.
| Crawler | Operator | robots.txt token | Blocked by |
|---|---|---|---|
| GPTBot
Crawls content that may be used to train OpenAI's generative AI foundation models. |
OpenAI | gptbot |
7%
2,145 of 30,177 |
| ClaudeBot
Collects web content that may contribute to training Anthropic's models; honors Crawl-delay. |
Anthropic | claudebot |
6%
1,743 of 30,177 |
| anthropic-ai
Legacy token widely blocked for Anthropic training; Anthropic now documents ClaudeBot, Claude-User and Claude-SearchBot. |
Anthropic | anthropic-ai |
4%
1,320 of 30,177 |
| Google-Extended
Control token with no user agent of its own; governs Gemini training and grounding use of Googlebot data. |
google-extendedno user agent † |
5%
1,391 of 30,177 |
|
| GoogleOther
Generic Google crawler used by product teams for one-off fetches such as internal research and development. |
googleother |
1%
383 of 30,177 |
|
| Applebot-Extended
Control token with no user agent; disallowing it excludes crawled content from Apple foundation model training. |
Apple | applebot-extendedno user agent † |
5%
1,378 of 30,177 |
| Meta-ExternalAgent
Crawls the web to train Meta's foundation AI models and to index content directly into products. |
Meta | meta-externalagent |
5%
1,430 of 30,177 |
| FacebookBot
Crawls public pages to improve language models behind Meta's speech recognition technology. |
Meta | facebookbot |
3%
979 of 30,177 |
| Meta-ExternalAds
Crawls the web to improve Meta's advertising and other business products and services. |
Meta | meta-externalads |
0.1%
22 of 30,177 |
| Bytespider ignores robots
Downloads content to train ByteDance LLMs and is widely reported to ignore robots.txt directives. |
ByteDance | bytespider |
6%
1,950 of 30,177 |
| Amazonbot
Crawls for Amazon product and Alexa answers and may use the content to train Amazon AI models. |
Amazon | amazonbot |
5%
1,411 of 30,177 |
| Diffbot
Extracts structured page data for Diffbot's knowledge graph, which is licensed to AI customers. |
Diffbot | diffbot |
4%
1,224 of 30,177 |
| omgili
Collects forum, news and blog content that Webz.io sells as web data feeds, including for AI training. |
Webz.io | omgili |
4%
1,171 of 30,177 |
| omgilibot
Legacy Omgili search crawler token still blocked alongside the current omgili agent. |
Webz.io | omgilibot |
4%
1,289 of 30,177 |
| AI2Bot
Collects web text for Ai2's open datasets used to train open language models such as OLMo. |
Allen Institute for AI | ai2bot |
2%
706 of 30,177 |
| cohere-training-data-crawler
Downloads training data for the large language models behind Cohere's enterprise AI products. |
Cohere | cohere-training-data-crawler |
2%
620 of 30,177 |
| MistralAI-Training
Crawls web content to build datasets for training Mistral's generative AI models. |
Mistral AI | mistralai-training |
0.5%
142 of 30,177 |
| PanguBot
Collects web content used to train Huawei's PanGu family of large models. |
Huawei | pangubot |
2%
518 of 30,177 |
| Timpibot
Crawls pages for Timpi's decentralized index, which is also used as LLM training data. |
Timpi | timpibot |
3%
833 of 30,177 |
| ImagesiftBot
Downloads public images plus surrounding text to build ImageSift's searchable image index. |
ImageSift (Hive) | imagesiftbot |
3%
853 of 30,177 |
| Kangaroo Bot
Scrapes site content into datasets used to train the Kangaroo LLM. |
Kangaroo LLM | kangaroo bot |
2%
466 of 30,177 |
| SemrushBot-OCOB
Crawls pages to feed Semrush's ContentShake AI writing tool. |
Semrush | semrushbot-ocob |
4%
1,126 of 30,177 |
| Scrapy
Generic scraping framework often used to build AI training datasets; obeys robots.txt only when ROBOTSTXT_OBEY is on. |
Zyte (open-source framework) | scrapy |
2%
745 of 30,177 |
† Google-Extended and Applebot-Extended are robots.txt control tokens with no user agent of their own. Nothing ever connects to a server identifying as either one: the fetching is done by Googlebot and Applebot, and the token only governs what the operator may do with what those crawlers already took. A live request test against these two is meaningless, so we measure them from the file alone.
Answer index 12 agents
These agents build the index an answer engine queries at question time. Disallowing one removes the site from those answers and from the citations that link back to it. Training is governed by separate tokens, so this is the expensive block, not the protective one.
| Crawler | Operator | robots.txt token | Blocked by |
|---|---|---|---|
| OAI-SearchBot
Indexes pages so they can be surfaced and cited in ChatGPT search results, not for training. |
OpenAI | oai-searchbot |
2%
650 of 30,177 |
| Claude-SearchBot
Indexes content to improve the relevance and accuracy of Claude's search results. |
Anthropic | claude-searchbot |
2%
484 of 30,177 |
| Googlebot
Crawls and renders pages for Google Search, Images, Video, News and Discover. |
googlebot |
0.1%
16 of 30,177 |
|
| Googlebot-News
Robots token controlling Google News inclusion; crawling itself uses the Googlebot user agents. |
googlebot-news |
0.1%
36 of 30,177 |
|
| Applebot
Crawls for Siri, Spotlight and Safari search; falls back to Googlebot rules and ignores Crawl-delay. |
Apple | applebot |
1%
343 of 30,177 |
| Bingbot
Indexes pages for Bing search and the Copilot answers that are grounded in the Bing index. |
Microsoft | bingbot |
0.1%
40 of 30,177 |
| msnbot
Legacy Microsoft search crawler token still honored alongside bingbot. |
Microsoft | msnbot |
0.1%
21 of 30,177 |
| PerplexityBot
Indexes and links pages in Perplexity search results; not used to collect foundation model training data. |
Perplexity | perplexitybot |
3%
1,054 of 30,177 |
| Meta-WebIndexer
Indexes pages so Meta AI can cite and link them in its search answers. |
Meta | meta-webindexer |
1%
359 of 30,177 |
| Amzn-SearchBot
Indexes content for Amazon search experiences such as Alexa; does not crawl for generative AI training. |
Amazon | amzn-searchbot |
0.8%
233 of 30,177 |
| MistralAI-Index
Indexes content for Mistral search behind Vibe answers; not used for generative AI training. |
Mistral AI | mistralai-index |
0.3%
81 of 30,177 |
| YouBot
Indexes pages for You.com search results and the AI answers built on that index. |
You.com | youbot |
3%
890 of 30,177 |
Live retrieval 12 agents
These agents fetch a single URL because a user asked about it. Disallowing one means the assistant answers from an older copy or from someone else's summary instead of the live page. Several are documented as not bound by robots.txt, because the request is user-initiated rather than a crawl.
| Crawler | Operator | robots.txt token | Blocked by |
|---|---|---|---|
| ChatGPT-User
Fetches a page when a ChatGPT user or GPT Action asks for it; user-initiated, so robots rules may not apply. |
OpenAI | chatgpt-user |
3%
1,053 of 30,177 |
| OAI-AdsBot
Visits pages submitted as ChatGPT ads to check policy compliance and ad relevance; not used for model training. |
OpenAI | oai-adsbot |
0.0%
11 of 30,177 |
| Claude-User
Retrieves pages on demand when a Claude user's question needs live web content. |
Anthropic | claude-user |
2%
498 of 30,177 |
| Google-CloudVertexBot
Crawls sites at a site owner's request to build Vertex AI agents; no effect on Google Search. |
google-cloudvertexbot |
1%
423 of 30,177 |
|
| Perplexity-User ignores robots
Fetches a page for a specific user question; Perplexity documents that it generally ignores robots.txt. |
Perplexity | perplexity-user |
2%
469 of 30,177 |
| Meta-ExternalFetcher ignores robots
Fetches individual links for agentic AI tasks; Meta documents that it may bypass robots.txt. |
Meta | meta-externalfetcher |
2%
557 of 30,177 |
| facebookexternalhit ignores robots
Fetches shared links for Facebook, Instagram and Messenger previews; may bypass robots.txt for integrity checks. |
Meta | facebookexternalhit |
0.2%
53 of 30,177 |
| TikTokSpider ignores robots
Fetches shared URLs for TikTok link previews and feeds; not expected to follow robots.txt. |
ByteDance | tiktokspider |
0.8%
255 of 30,177 |
| Amzn-User ignores robots
Fetches live pages to answer a user's Alexa question; Amazon documents it may not follow all robots.txt rules. |
Amazon | amzn-user |
0.6%
167 of 30,177 |
| cohere-ai
Retrieves pages to answer user-initiated prompts in Cohere's enterprise AI products. |
Cohere | cohere-ai |
4%
1,223 of 30,177 |
| MistralAI-User
Fetches pages on demand so Mistral's Vibe assistant can answer a question with live, cited web content. |
Mistral AI | mistralai-user |
1%
447 of 30,177 |
| DuckAssistBot
Crawls pages in real time for DuckDuckGo's cited AI-assisted answers; not used for model training. |
DuckDuckGo | duckassistbot |
2%
519 of 30,177 |
Web archive 1 agents
These agents build public web archives. Those archives are then used as pretraining corpora by parties who never crawled the site themselves, so one disallow here reaches model builders that never appear in a log file.
| Crawler | Operator | robots.txt token | Blocked by |
|---|---|---|---|
| CCBot
Builds the open Common Crawl web archive, a common source of LLM pretraining corpora. |
Common Crawl Foundation | ccbot |
7%
2,066 of 30,177 |
How to read the token column reference
The token is what robots.txt matches on, case-insensitively, against the longest matching product name in the incoming User-Agent header. Writing User-agent: gptbot creates a group that applies to GPTBot and to nothing else, and that group replaces the * group for GPTBot completely. Any disallow you rely on in the wildcard group has to be repeated in the specific one, which is the single most common way a site accidentally opens a path it meant to close.
A group named for an agent that never visits costs nothing, and a group named for an agent that does not honor the file buys nothing. The ignores robots marker above records where the operator itself documents that the agent may fetch regardless, which moves the decision from robots.txt to the edge.
Measure your own site free
The registry says what each agent is. A scan says which of them your site currently admits, including the ones your robots.txt permits and your edge refuses.