tenjin.io
Open to answer engines, publishes llms.txt.
Whether AI crawlers and answer-engine fetchers are permitted to request the page at all, in robots.txt, in robots directives, and at the edge.
Whether a fetcher that does not execute JavaScript receives the actual content, in markup an extractor can segment.
Machine-readable markup that states the page's type, entities, canonical URL, and discrete facts instead of leaving them to be inferred.
Signals that let an answer engine name the author, date the content, resolve the publisher, and cite it under known terms.
Who is allowed to read this site 0 of 24 answer engines blocked
| Crawler | Operator | Uses content for | robots.txt | Live request |
|---|---|---|---|---|
| GPTBot Crawls content that may be used to train OpenAI's generative AI foundation models. |
OpenAI | Model training | allowed Allow: / |
served 200 |
| OAI-SearchBot Indexes pages so they can be surfaced and cited in ChatGPT search results, not for training. |
OpenAI | Answer index | allowed Allow: / |
served 200 |
| ChatGPT-User Fetches a page when a ChatGPT user or GPT Action asks for it; user-initiated, so robots rules may not apply. |
OpenAI | Live retrieval | allowed Allow: / |
not probed |
| OAI-AdsBot Visits pages submitted as ChatGPT ads to check policy compliance and ad relevance; not used for model training. |
OpenAI | Live retrieval | allowed Allow: / |
not probed |
| ClaudeBot Collects web content that may contribute to training Anthropic's models; honors Crawl-delay. |
Anthropic | Model training | allowed Allow: / |
served 200 |
| Claude-User Retrieves pages on demand when a Claude user's question needs live web content. |
Anthropic | Live retrieval | allowed Allow: / |
not probed |
| Claude-SearchBot Indexes content to improve the relevance and accuracy of Claude's search results. |
Anthropic | Answer index | allowed Allow: / |
not probed |
| anthropic-ai Legacy token widely blocked for Anthropic training; Anthropic now documents ClaudeBot, Claude-User and Claude-SearchBot. |
Anthropic | Model training | allowed Allow: / |
not probed |
| Google-Extended Control token with no user agent of its own; governs Gemini training and grounding use of Googlebot data. |
Model training | allowed Allow: / |
not probed | |
| Googlebot Crawls and renders pages for Google Search, Images, Video, News and Discover. |
Answer index | allowed Allow: / |
not probed | |
| Googlebot-News Robots token controlling Google News inclusion; crawling itself uses the Googlebot user agents. |
Answer index | allowed Allow: / |
not probed | |
| Google-CloudVertexBot Crawls sites at a site owner's request to build Vertex AI agents; no effect on Google Search. |
Live retrieval | allowed Allow: / |
not probed | |
| GoogleOther Generic Google crawler used by product teams for one-off fetches such as internal research and development. |
Model training | allowed Allow: / |
not probed | |
| Applebot Crawls for Siri, Spotlight and Safari search; falls back to Googlebot rules and ignores Crawl-delay. |
Apple | Answer index | allowed Allow: / |
not probed |
| Applebot-Extended Control token with no user agent; disallowing it excludes crawled content from Apple foundation model training. |
Apple | Model training | allowed Allow: / |
not probed |
| Bingbot Indexes pages for Bing search and the Copilot answers that are grounded in the Bing index. |
Microsoft | Answer index | allowed Allow: / |
not probed |
| msnbot Legacy Microsoft search crawler token still honored alongside bingbot. |
Microsoft | Answer index | allowed Allow: / |
not probed |
| PerplexityBot Indexes and links pages in Perplexity search results; not used to collect foundation model training data. |
Perplexity | Answer index | allowed Allow: / |
served 200 |
| Perplexity-User ignores robots Fetches a page for a specific user question; Perplexity documents that it generally ignores robots.txt. |
Perplexity | Live retrieval | allowed Allow: / |
not probed |
| Meta-ExternalAgent Crawls the web to train Meta's foundation AI models and to index content directly into products. |
Meta | Model training | allowed Allow: / |
not probed |
| Meta-ExternalFetcher ignores robots Fetches individual links for agentic AI tasks; Meta documents that it may bypass robots.txt. |
Meta | Live retrieval | allowed Allow: / |
not probed |
| FacebookBot Crawls public pages to improve language models behind Meta's speech recognition technology. |
Meta | Model training | allowed Allow: / |
not probed |
| Meta-WebIndexer Indexes pages so Meta AI can cite and link them in its search answers. |
Meta | Answer index | allowed Allow: / |
not probed |
| Meta-ExternalAds Crawls the web to improve Meta's advertising and other business products and services. |
Meta | Model training | allowed Allow: / |
not probed |
| facebookexternalhit ignores robots Fetches shared links for Facebook, Instagram and Messenger previews; may bypass robots.txt for integrity checks. |
Meta | Live retrieval | allowed Allow: / |
not probed |
| Bytespider ignores robots Downloads content to train ByteDance LLMs and is widely reported to ignore robots.txt directives. |
ByteDance | Model training | allowed Allow: / |
not probed |
| TikTokSpider ignores robots Fetches shared URLs for TikTok link previews and feeds; not expected to follow robots.txt. |
ByteDance | Live retrieval | allowed Allow: / |
not probed |
| Amazonbot Crawls for Amazon product and Alexa answers and may use the content to train Amazon AI models. |
Amazon | Model training | allowed Allow: / |
not probed |
| Amzn-SearchBot Indexes content for Amazon search experiences such as Alexa; does not crawl for generative AI training. |
Amazon | Answer index | allowed Allow: / |
not probed |
| Amzn-User ignores robots Fetches live pages to answer a user's Alexa question; Amazon documents it may not follow all robots.txt rules. |
Amazon | Live retrieval | allowed Allow: / |
not probed |
| CCBot Builds the open Common Crawl web archive, a common source of LLM pretraining corpora. |
Common Crawl Foundation | Archive | allowed Allow: / |
not probed |
| Diffbot Extracts structured page data for Diffbot's knowledge graph, which is licensed to AI customers. |
Diffbot | Model training | allowed Allow: / |
not probed |
| omgili Collects forum, news and blog content that Webz.io sells as web data feeds, including for AI training. |
Webz.io | Model training | allowed Allow: / |
not probed |
| omgilibot Legacy Omgili search crawler token still blocked alongside the current omgili agent. |
Webz.io | Model training | allowed Allow: / |
not probed |
| AI2Bot Collects web text for Ai2's open datasets used to train open language models such as OLMo. |
Allen Institute for AI | Model training | allowed Allow: / |
not probed |
| cohere-ai Retrieves pages to answer user-initiated prompts in Cohere's enterprise AI products. |
Cohere | Live retrieval | allowed Allow: / |
not probed |
| cohere-training-data-crawler Downloads training data for the large language models behind Cohere's enterprise AI products. |
Cohere | Model training | allowed Allow: / |
not probed |
| MistralAI-User Fetches pages on demand so Mistral's Vibe assistant can answer a question with live, cited web content. |
Mistral AI | Live retrieval | allowed Allow: / |
not probed |
| MistralAI-Index Indexes content for Mistral search behind Vibe answers; not used for generative AI training. |
Mistral AI | Answer index | allowed Allow: / |
not probed |
| MistralAI-Training Crawls web content to build datasets for training Mistral's generative AI models. |
Mistral AI | Model training | allowed Allow: / |
not probed |
| DuckAssistBot Crawls pages in real time for DuckDuckGo's cited AI-assisted answers; not used for model training. |
DuckDuckGo | Live retrieval | allowed Allow: / |
not probed |
| YouBot Indexes pages for You.com search results and the AI answers built on that index. |
You.com | Answer index | allowed Allow: / |
not probed |
| PanguBot Collects web content used to train Huawei's PanGu family of large models. |
Huawei | Model training | allowed Allow: / |
not probed |
| Timpibot Crawls pages for Timpi's decentralized index, which is also used as LLM training data. |
Timpi | Model training | allowed Allow: / |
not probed |
| ImagesiftBot Downloads public images plus surrounding text to build ImageSift's searchable image index. |
ImageSift (Hive) | Model training | allowed Allow: / |
not probed |
| Kangaroo Bot Scrapes site content into datasets used to train the Kangaroo LLM. |
Kangaroo LLM | Model training | allowed Allow: / |
not probed |
| SemrushBot-OCOB Crawls pages to feed Semrush's ContentShake AI writing tool. |
Semrush | Model training | allowed Allow: / |
not probed |
| Scrapy Generic scraping framework often used to build AI training datasets; obeys robots.txt only when ROBOTSTXT_OBEY is on. |
Zyte (open-source framework) | Model training | allowed Allow: / |
not probed |
Reach 40 / 40
Answer-engine fetchers are allowed to retrieve and cite this page
All 24 answer-engine fetchers are allowed to retrieve pages for citation.
Training crawlers may fetch this path
All 23 tracked training crawlers are allowed.
AI user agents receive the same 200 response as browsers
Live requests as 4 AI user agents were served normally.
No blanket disallow applies to this path
The wildcard group does not disallow the entire site.
robots.txt served as plain text with a 200 response
robots.txt served, 2234 bytes, 15 group(s).
No Crawl-delay directive constrains fetchers
No Crawl-delay directive.
Page is indexable, with no noindex directive
No noindex directive on the homepage.
Full-length snippet extraction is permitted
Snippets are not restricted by meta tags.
X-Robots-Tag header is absent or permissive
No restrictive X-Robots-Tag header.
Readability 22.7 / 25
Visible text is a small fraction of the HTML payload
Text is 3.4% of the 312 KB document; 69 KB is inline script.
Fix. Move inline hydration state and large inline scripts out of the document, or fetch them after load instead of embedding them. Flatten wrapper `div` trees and let semantic elements carry the content, and keep utility-class soup out of the article body. Serve the same text without the boilerplate at a stable URL if you need a clean extraction target.
Substantive text is present in the server-rendered HTML
1644 words of text are present in the raw HTML (wordpress). Most AI fetchers do not run JavaScript.
Primary content is wrapped in a semantic landmark
An <article> element marks the primary content.
Headings form a single, ordered outline
1 H1 and 22 headings total.
Title is unique and describes the page in specific terms
Title is 35 characters: "Mobile Marketing Analytics | Tenjin"
Meta description provides an author-written summary
Meta description is 144 characters.
Document language is declared on the html element
Declared language: en-GB.
Structure 20 / 20
Page ships JSON-LD structured data
10 JSON-LD node(s): WebPage, ReadAction, BreadcrumbList, ListItem, WebSite, SearchAction.
Structured data uses a specific type that matches the page
Recognized types: breadcrumblist, website, organization.
JSON-LD parses cleanly with recognised schema.org terms
All JSON-LD blocks parse cleanly.
Page declares a self-referential canonical URL
Canonical: https://tenjin.com/
XML sitemap is declared in robots.txt and returns 200
Sitemap found at /sitemap.xml (5 URLs on the first document).
Key facts are available in lists or tables
0 tables, 33 lists, 0 code blocks, 0 question headings.
Attribution 11 / 15
No machine-readable author is attached to the page
No author or Person entity, which weakens the authority signals answer engines use.
Fix. Add an `author` property to the page's `Article`, `BlogPosting`, or `NewsArticle` node, typed as `Person` or `Organization`, with a `name` and a `url` pointing at a real profile page. Give each author a stable `@id` and reuse it across posts so the entity consolidates. Keep the visible byline identical to the structured value, and avoid generic names such as "Admin" or "Staff Writer" where a real author exists.
{
"@context": "https://schema.org",
"@type": "BlogPosting",
"headline": "Allow AI answer engines in robots.txt",
"author": {
"@type": "Person",
"@id": "https://example.com/authors/dana-reyes#person",
"name": "Dana Reyes",
"url": "https://example.com/authors/dana-reyes",
"jobTitle": "Infrastructure Engineer",
"sameAs": ["https://github.com/danareyes"]
}
}ReferenceNo machine-readable license or usage terms for the content
No licence declaration, so reuse terms are ambiguous.
Fix. Add a `license` property to the page's structured data pointing at a specific license URL, such as a Creative Commons deed or your own terms page, and add `rel="license"` on the visible link. Use `usageInfo` for conditions that are not a standard license, such as attribution wording or an API-only clause. State the terms once, at a stable URL, and reference it from every page rather than restating it per template.
<a rel="license" href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a>
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "Allow AI answer engines in robots.txt",
"license": "https://creativecommons.org/licenses/by/4.0/",
"usageInfo": "https://example.com/legal/content-reuse",
"creditText": "Example, crawlcensus research desk"
}
</script>Reference/llms.txt lists canonical pages in the documented format
/llms.txt is present and well formed (6 sections, 25 links).
Published and modified dates are declared in ISO 8601
dateModified is published.
Organization entity declares the publisher and its identifiers
Organization entity with sameAs links is present.
Evidence raw measurements
- Size
- 2,234 bytes
- Groups
- 15
- Sitemaps
- https://tenjin.com/sitemap_index.xml
View the file as our crawler received it
# START YOAST BLOCK
# ---------------------------
User-agent: *
Allow: /
Sitemap: https://tenjin.com/sitemap_index.xml
Allow: /wp-admin/admin-ajax.php
Disallow: /wp-admin/
Disallow:/*?__hstc
Disallow:/vi/blog/
Disallow:/blog/author/
Disallow:/zh/blog/author/
Disallow:/ja/blog/author/
Disallow:/ru/blog/author/
Disallow:/blog/tag/
Disallow:/zh/blog/tag/
Disallow:/ja/blog/tag/
Disallow:/ru/blog/tag/
Disallow:/integration/
Disallow:/zh/integration/
Disallow:/ja/integration/
Disallow:/ru/integration/
Disallow:/zh/blog/category/
Disallow:/ja/blog/category/
Disallow:/ru/blog/category/
Disallow:/blog/page/
Disallow:/ja/blog/page/
Disallow:/ru/blog/page/
Disallow:/zh/blog/page/
Disallow:/vi/privacy/
Disallow:/es/privacy/
Disallow:/vi/disclosure/
Disallow:/es/tenjin-inc-data-processing-addendum/
Disallow:/ru/privacy/
Disallow:/ru/resources/
Disallow:/pt/privacy/
Disallow:/ru/tenjin-inc-data-processing-addendum/
Disallow:/es/terms/
Disallow:/pt/tenjin-inc-data-processing-addendum/
Disallow:/ja/terms/
Disallow:/vi/tenjin-inc-data-processing-addendum/
Disallow:/es/tenjin-inc-data-processing-addendum/
Disallow:/es/docs/
Disallow:/vi/docs/
Disallow:/ru/docs/
Disallow:/pt/docs/
Disallow:/es/glossary/
Disallow:/vi/glossary/
Disallow:/ru/glossary/
Disallow:/pt/glossary/
Disallow:/pt/resources/
Disallow:/vi/resources/
Disallow:/ru/resources/
Disallow:/zh/resources/
Disallow:/ja/resources/
Disallow:/es/resources/
Disallow:/resources/
# --- NEW AI BOTS ALLOW LIST ---
# OpenAI (ChatGPT & Search)
User-agent: GPTBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: OAI-SearchBot
Allow: /
# Google AI
User-agent: Google-Extended
Allow: /
# Perplexity
User-agent: PerplexityBot
Allow: /
# Anthropic (Claude)
User-agent: ClaudeBot
Allow: /
# Apple Intelligence
User-agent: Applebot-Extended
Allow: /
# Meta AI
User-agent: Meta-ExternalAgent
Allow: /
User-agent: Meta-ExternalFetcher
Allow: /
# Foundational Data & Other AI
User-agent: CCBot
Allow: /
User-agent: Amazonbot
Allow: /
User-agent: cohere-ai
Allow: /
User-agent: Omgilibot
Allow: /
User-agent: Omgili
Allow: /
# ---------------------------
# END YOAST BLOCK- llms.txt
- valid, 2,538 bytes, 25 links
- ai.txt
- absent
- Sitemap
- /sitemap.xml (5 URLs)
- Feeds
- https://tenjin.com/feed/
https://tenjin.com/comments/feed/ - Schema types
- WebPage, ReadAction, BreadcrumbList, ListItem, WebSite, SearchAction, EntryPoint, PropertyValueSpecification, Organization, ImageObject
- Edge
- Cloudflare
- Final URL
- https://tenjin.com/
- HTML size
- 312 KB, 1,644 words of text
Publish the score free
Sites that score well embed the badge. It links back to this live report, which re-measures on every scan.
<a href="https://crawlcensus.com/site/tenjin.io"><img src="https://crawlcensus.com/badge/tenjin.io.svg" alt="AI access score for tenjin.io" width="150" height="20"></a>Track changes on this domain pro
Crawler policy is edited quietly. We re-scan monitored domains daily, keep the history, and email you the moment a crawler is blocked or unblocked, an llms.txt appears, or the score moves.
One address, unlimited domains during the beta. No newsletter, only change alerts.