yhprn.com
Open to answer engines, refuses AI user agents at the edge, no structured data.
Whether AI crawlers and answer-engine fetchers are permitted to request the page at all. What happens on the wire outweighs what the policy file says: a live refusal costs 16 of these 40 points, twice what the heaviest robots.txt rule costs, because a page a crawler cannot fetch is unreadable no matter how well it is marked up.
Whether a fetcher that does not execute JavaScript receives the actual content, in markup an extractor can segment.
Machine-readable markup that states the page's type, entities, canonical URL, and discrete facts instead of leaving them to be inferred.
Signals that let an answer engine name the author, date the content, resolve the publisher, and cite it under known terms.
Who is allowed to read this site 0 of 24 answer engines blocked
| Crawler | Operator | Uses content for | robots.txt | Live request |
|---|---|---|---|---|
| GPTBot Crawls content that may be used to train OpenAI's generative AI foundation models. |
OpenAI | Model training | allowed | refused 403 |
| OAI-SearchBot Indexes pages so they can be surfaced and cited in ChatGPT search results, not for training. |
OpenAI | Answer index | allowed | refused 403 |
| ChatGPT-User Fetches a page when a ChatGPT user or GPT Action asks for it; user-initiated, so robots rules may not apply. |
OpenAI | Live retrieval | allowed | not probed |
| OAI-AdsBot Visits pages submitted as ChatGPT ads to check policy compliance and ad relevance; not used for model training. |
OpenAI | Live retrieval | allowed | not probed |
| ClaudeBot Collects web content that may contribute to training Anthropic's models; honors Crawl-delay. |
Anthropic | Model training | allowed | refused 403 |
| Claude-User Retrieves pages on demand when a Claude user's question needs live web content. |
Anthropic | Live retrieval | allowed | not probed |
| Claude-SearchBot Indexes content to improve the relevance and accuracy of Claude's search results. |
Anthropic | Answer index | allowed | not probed |
| anthropic-ai Legacy token widely blocked for Anthropic training; Anthropic now documents ClaudeBot, Claude-User and Claude-SearchBot. |
Anthropic | Model training | allowed | not probed |
| Google-Extended Control token with no user agent of its own; governs Gemini training and grounding use of Googlebot data. |
Model training | allowed | not probed | |
| Googlebot Crawls and renders pages for Google Search, Images, Video, News and Discover. |
Answer index | allowed | not probed | |
| Googlebot-News Robots token controlling Google News inclusion; crawling itself uses the Googlebot user agents. |
Answer index | allowed | not probed | |
| Google-CloudVertexBot Crawls sites at a site owner's request to build Vertex AI agents; no effect on Google Search. |
Live retrieval | allowed | not probed | |
| GoogleOther Generic Google crawler used by product teams for one-off fetches such as internal research and development. |
Model training | allowed | not probed | |
| Applebot Crawls for Siri, Spotlight and Safari search; falls back to Googlebot rules and ignores Crawl-delay. |
Apple | Answer index | allowed | not probed |
| Applebot-Extended Control token with no user agent; disallowing it excludes crawled content from Apple foundation model training. |
Apple | Model training | allowed | not probed |
| Bingbot Indexes pages for Bing search and the Copilot answers that are grounded in the Bing index. |
Microsoft | Answer index | allowed | not probed |
| msnbot Legacy Microsoft search crawler token still honored alongside bingbot. |
Microsoft | Answer index | allowed | not probed |
| PerplexityBot Indexes and links pages in Perplexity search results; not used to collect foundation model training data. |
Perplexity | Answer index | allowed | refused 403 |
| Perplexity-User ignores robots Fetches a page for a specific user question; Perplexity documents that it generally ignores robots.txt. |
Perplexity | Live retrieval | allowed | not probed |
| Meta-ExternalAgent Crawls the web to train Meta's foundation AI models and to index content directly into products. |
Meta | Model training | allowed | not probed |
| Meta-ExternalFetcher ignores robots Fetches individual links for agentic AI tasks; Meta documents that it may bypass robots.txt. |
Meta | Live retrieval | allowed | not probed |
| FacebookBot Crawls public pages to improve language models behind Meta's speech recognition technology. |
Meta | Model training | allowed | not probed |
| Meta-WebIndexer Indexes pages so Meta AI can cite and link them in its search answers. |
Meta | Answer index | allowed | not probed |
| Meta-ExternalAds Crawls the web to improve Meta's advertising and other business products and services. |
Meta | Model training | allowed | not probed |
| facebookexternalhit ignores robots Fetches shared links for Facebook, Instagram and Messenger previews; may bypass robots.txt for integrity checks. |
Meta | Live retrieval | allowed | not probed |
| Bytespider ignores robots Downloads content to train ByteDance LLMs and is widely reported to ignore robots.txt directives. |
ByteDance | Model training | allowed | not probed |
| TikTokSpider ignores robots Fetches shared URLs for TikTok link previews and feeds; not expected to follow robots.txt. |
ByteDance | Live retrieval | allowed | not probed |
| Amazonbot Crawls for Amazon product and Alexa answers and may use the content to train Amazon AI models. |
Amazon | Model training | allowed | not probed |
| Amzn-SearchBot Indexes content for Amazon search experiences such as Alexa; does not crawl for generative AI training. |
Amazon | Answer index | allowed | not probed |
| Amzn-User ignores robots Fetches live pages to answer a user's Alexa question; Amazon documents it may not follow all robots.txt rules. |
Amazon | Live retrieval | allowed | not probed |
| CCBot Builds the open Common Crawl web archive, a common source of LLM pretraining corpora. |
Common Crawl Foundation | Archive | allowed | not probed |
| Diffbot Extracts structured page data for Diffbot's knowledge graph, which is licensed to AI customers. |
Diffbot | Model training | allowed | not probed |
| omgili Collects forum, news and blog content that Webz.io sells as web data feeds, including for AI training. |
Webz.io | Model training | allowed | not probed |
| omgilibot Legacy Omgili search crawler token still blocked alongside the current omgili agent. |
Webz.io | Model training | allowed | not probed |
| AI2Bot Collects web text for Ai2's open datasets used to train open language models such as OLMo. |
Allen Institute for AI | Model training | allowed | not probed |
| cohere-ai Retrieves pages to answer user-initiated prompts in Cohere's enterprise AI products. |
Cohere | Live retrieval | allowed | not probed |
| cohere-training-data-crawler Downloads training data for the large language models behind Cohere's enterprise AI products. |
Cohere | Model training | allowed | not probed |
| MistralAI-User Fetches pages on demand so Mistral's Vibe assistant can answer a question with live, cited web content. |
Mistral AI | Live retrieval | allowed | not probed |
| MistralAI-Index Indexes content for Mistral search behind Vibe answers; not used for generative AI training. |
Mistral AI | Answer index | allowed | not probed |
| MistralAI-Training Crawls web content to build datasets for training Mistral's generative AI models. |
Mistral AI | Model training | allowed | not probed |
| DuckAssistBot Crawls pages in real time for DuckDuckGo's cited AI-assisted answers; not used for model training. |
DuckDuckGo | Live retrieval | allowed | not probed |
| YouBot Indexes pages for You.com search results and the AI answers built on that index. |
You.com | Answer index | allowed | not probed |
| PanguBot Collects web content used to train Huawei's PanGu family of large models. |
Huawei | Model training | allowed | not probed |
| Timpibot Crawls pages for Timpi's decentralized index, which is also used as LLM training data. |
Timpi | Model training | allowed | not probed |
| ImagesiftBot Downloads public images plus surrounding text to build ImageSift's searchable image index. |
ImageSift (Hive) | Model training | allowed | not probed |
| Kangaroo Bot Scrapes site content into datasets used to train the Kangaroo LLM. |
Kangaroo LLM | Model training | allowed | not probed |
| SemrushBot-OCOB Crawls pages to feed Semrush's ContentShake AI writing tool. |
Semrush | Model training | allowed | not probed |
| Scrapy Generic scraping framework often used to build AI training datasets; obeys robots.txt only when ROBOTSTXT_OBEY is on. |
Zyte (open-source framework) | Model training | allowed | not probed |
Reach 23.5 / 40
Edge returns 403, 429, or a challenge to AI user agents
Requests identifying as gptbot, oai-searchbot, perplexitybot, claudebot were refused at the edge (HTTP 403 Forbidden; HTTP 403 Forbidden; HTTP 403 Forbidden; Cloudflare challenge (cf-mitigated)).
Fix. Fetch the page with each AI user agent string and compare the status and byte count against a browser request. On Cloudflare, check whether the "Block AI bots" toggle in AI Crawl Control, a Bot Fight Mode rule, or a WAF custom rule on `cf.verified_bot_category` is catching the request, then narrow it: block the training category and add a skip rule for the retrieval agents you want citing you. Verified bots must not be handed Managed Challenge, since a challenge is a hard failure for a non-browser client. Re-test after every WAF or bot-management change, because these toggles are applied zone-wide.
for ua in "OAI-SearchBot/1.0" "ChatGPT-User/1.0" "Claude-User/1.0" \
"Claude-SearchBot/1.0" "PerplexityBot/1.0" "Mozilla/5.0"; do
code=$(curl -s -o /dev/null -w '%{http_code}' -A "$ua" https://example.com/)
printf '%s\t%s\n' "$code" "$ua"
doneReferencerobots.txt sets Crawl-delay, which most fetchers ignore
Crawl-delay of 1s is declared. Large crawlers may fetch only a few pages per hour, or ignore the directive entirely.
Fix. Remove `Crawl-delay` and handle load at the edge instead, with caching and rate limiting keyed on the client. If crawl volume is the real problem, cache HTML at your CDN so repeat fetches never reach the origin. Keep the file to standard directives (`User-agent`, `Allow`, `Disallow`, `Sitemap`) so parser behaviour is predictable.
# Remove non-standard throttling directives:
# Crawl-delay: 10
User-agent: *
Allow: /ReferenceAnswer-engine fetchers are allowed to retrieve and cite this page
All 24 answer-engine fetchers are allowed to retrieve pages for citation.
Training crawlers may fetch this path
All 23 tracked training crawlers are allowed.
No blanket disallow applies to this path
The wildcard group does not disallow the entire site.
robots.txt served as plain text with a 200 response
robots.txt served, 7464 bytes, 5 group(s).
Page is indexable, with no noindex directive
No noindex directive on the homepage.
Full-length snippet extraction is permitted
Snippets are not restricted by meta tags.
X-Robots-Tag header is absent or permissive
No restrictive X-Robots-Tag header.
Readability 21.5 / 25
Visible text is a small fraction of the HTML payload
Text is 4.0% of the 113 KB document; 6 KB is inline script.
Fix. Move inline hydration state and large inline scripts out of the document, or fetch them after load instead of embedding them. Flatten wrapper `div` trees and let semantic elements carry the content, and keep utility-class soup out of the article body. Serve the same text without the boilerplate at a stable URL if you need a clean extraction target.
Heading levels are missing, duplicated, or skipped
2 H1 and 3 headings total.
Fix. Give every page one `h1` that names its subject, then nest `h2` and `h3` without skipping levels. Make each heading describe the section beneath it in words a reader would search for, rather than a label like "Overview". Style with CSS instead of choosing heading levels for their font size, and never use a heading tag for a caption or a button.
<h1>Robots.txt rules for AI crawlers</h1>
<h2>Training crawlers</h2>
<h3>GPTBot</h3>
<h3>ClaudeBot</h3>
<h2>Retrieval and citation fetchers</h2>
<h3>OAI-SearchBot</h3>ReferenceSubstantive text is present in the server-rendered HTML
781 words of text are present in the raw HTML. Most AI fetchers do not run JavaScript.
Primary content is wrapped in a semantic landmark
An <article> element marks the primary content.
Title is unique and describes the page in specific terms
Title is 53 characters: "Free Prn Movies & HD XXX Sex Videos, Adult Prno Films"
Meta description provides an author-written summary
Meta description is 117 characters.
Document language is declared on the html element
Declared language: en.
Structure 8.5 / 20
No JSON-LD structured data found on the page
No JSON-LD structured data on the homepage.
Fix. Add one `<script type="application/ld+json">` block describing the page's primary entity, using the schema.org type that actually fits: `Article` or `NewsArticle`, `Product`, `Recipe`, `Event`, `FAQPage`, or `SoftwareApplication`. Populate the required properties for that type and make every value match visible page content. Prefer JSON-LD over microdata or RDFa, since it is the format the major crawlers document, and render it server-side so non-JavaScript fetchers see it.
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "Allow AI answer engines in robots.txt",
"description": "How to opt out of model training while staying citable.",
"url": "https://example.com/blog/robots-for-ai",
"mainEntityOfPage": {
"@type": "WebPage",
"@id": "https://example.com/blog/robots-for-ai"
},
"image": "https://example.com/images/robots-for-ai.png",
"inLanguage": "en",
"datePublished": "2026-02-11T09:00:00-05:00",
"dateModified": "2026-08-04T14:20:00-04:00",
"author": {
"@type": "Person",
"name": "Dana Reyes",
"url": "https://example.com/authors/dana-reyes"
},
"publisher": {
"@type": "Organization",
"name": "Example",
"url": "https://example.com",
"logo": {
"@type": "ImageObject",
"url": "https://example.com/logo.png",
"width": 512,
"height": 512
}
}
}ReferenceStructured data is too generic for what the page is about
No entity types that answer engines consume.
Fix. Replace bare `WebPage` and `WebSite` nodes with the most specific type that describes the page, and fill the properties that type defines. Use `@graph` to publish several linked nodes on one page, such as an `Article` whose `publisher` points at an `Organization` node by `@id`. Add `BreadcrumbList` for hierarchy and reuse the same `@id` values across pages so the entity resolves to one record.
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "Product",
"@id": "https://example.com/products/widget#product",
"name": "Widget Pro",
"sku": "WGT-PRO-1",
"brand": { "@type": "Brand", "name": "Example" },
"offers": {
"@type": "Offer",
"url": "https://example.com/products/widget",
"price": "49.00",
"priceCurrency": "USD",
"availability": "https://schema.org/InStock"
}
},
{
"@type": "BreadcrumbList",
"itemListElement": [
{ "@type": "ListItem", "position": 1, "name": "Products", "item": "https://example.com/products" },
{ "@type": "ListItem", "position": 2, "name": "Widget Pro" }
]
}
]
}ReferenceFacts are only in prose, with no list or table structure
0 tables, 2 lists, 0 code blocks, 0 question headings.
Fix. Put specifications, comparisons, pricing, and steps into real `<table>` markup with `<caption>` and `<th scope>`, or into `<ul>`, `<ol>`, and `<dl>` elements. Do not simulate tables with `div` grids, and avoid images of tables, which carry no extractable text. Keep one fact per row or list item so a chunk stays meaningful on its own.
<table>
<caption>AI fetcher purposes</caption>
<thead>
<tr><th scope="col">User agent</th><th scope="col">Purpose</th></tr>
</thead>
<tbody>
<tr><td>GPTBot</td><td>Model training</td></tr>
<tr><td>OAI-SearchBot</td><td>Retrieval and citation</td></tr>
</tbody>
</table>ReferenceJSON-LD parses cleanly with recognised schema.org terms
All JSON-LD blocks parse cleanly.
Page declares a self-referential canonical URL
Canonical: https://www.yhprn.com/
XML sitemap is declared in robots.txt and returns 200
Sitemap found at /sitemap.xml (5 URLs on the first document).
Attribution 1 / 9
No /llms.txt index of canonical pages
No /llms.txt.
Fix. Publish `/llms.txt` as `text/plain` markdown: an `#` H1 with the project name, a `>` blockquote summary, optional plain paragraphs of context, then `##` sections whose bullets are `[title](absolute-url): note`. Link the pages you want quoted, put lower-priority links under an `## Optional` section, and prefer URLs that also serve clean markdown. Keep it generated from the same source as your sitemap so it does not drift, and remember it is a hint for assistants, not an access control mechanism.
# Example
> Example publishes reference documentation for the Widget API and guides for
> configuring crawler access.
Prefer the pages below over search results; each URL is canonical.
## Docs
- [Widget API reference](https://example.com/docs/api): endpoints, auth, limits.
- [Quickstart](https://example.com/docs/quickstart): first request in five minutes.
## Policies
- [Crawler policy](https://example.com/legal/crawlers): which agents we allow.
## Optional
- [Changelog](https://example.com/changelog): dated release notes.ReferenceNo Organization entity identifies the publisher
No Organization entity, so the brand is harder to resolve to a known entity.
Fix. Publish one `Organization` node, usually on the home page, with `name`, `url`, `logo`, and `sameAs` pointing at the profiles that already describe you: Wikipedia or Wikidata, Crunchbase, LinkedIn, and your primary social accounts. Give it a stable `@id` such as `https://example.com/#organization` and reference that `@id` from each page's `publisher` property instead of repeating the block. Add `contactPoint` and `address` when they are public, and keep every value identical to what the site shows.
{
"@context": "https://schema.org",
"@type": "Organization",
"@id": "https://example.com/#organization",
"name": "Example",
"legalName": "Example Holdings, Inc.",
"url": "https://example.com",
"logo": {
"@type": "ImageObject",
"url": "https://example.com/logo.png",
"width": 512,
"height": 512
},
"sameAs": [
"https://www.wikidata.org/wiki/Q00000000",
"https://www.linkedin.com/company/example",
"https://github.com/example"
],
"contactPoint": {
"@type": "ContactPoint",
"contactType": "customer support",
"email": "support@example.com",
"areaServed": "US",
"availableLanguage": ["en"]
}
}ReferenceNo machine-readable license or usage terms for the content
No licence declaration, so reuse terms are ambiguous.
Fix. Add a `license` property to the page's structured data pointing at a specific license URL, such as a Creative Commons deed or your own terms page, and add `rel="license"` on the visible link. Use `usageInfo` for conditions that are not a standard license, such as attribution wording or an API-only clause. State the terms once, at a stable URL, and reference it from every page rather than restating it per template.
<a rel="license" href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a>
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "Allow AI answer engines in robots.txt",
"license": "https://creativecommons.org/licenses/by/4.0/",
"usageInfo": "https://example.com/legal/content-reuse",
"creditText": "Example, crawlcensus research desk"
}
</script>ReferenceNo machine-readable author is attached to the page
This page declares no article-type structured data, so a machine-readable byline is not expected of it. The check is excluded from the score rather than counted as a failure.
Fix. Add an `author` property to the page's `Article`, `BlogPosting`, or `NewsArticle` node, typed as `Person` or `Organization`, with a `name` and a `url` pointing at a real profile page. Give each author a stable `@id` and reuse it across posts so the entity consolidates. Keep the visible byline identical to the structured value, and avoid generic names such as "Admin" or "Staff Writer" where a real author exists.
{
"@context": "https://schema.org",
"@type": "BlogPosting",
"headline": "Allow AI answer engines in robots.txt",
"author": {
"@type": "Person",
"@id": "https://example.com/authors/dana-reyes#person",
"name": "Dana Reyes",
"url": "https://example.com/authors/dana-reyes",
"jobTitle": "Infrastructure Engineer",
"sameAs": ["https://github.com/danareyes"]
}
}ReferenceNo machine-readable published or modified date
This page declares no article-type structured data, so publication and modification dates are not expected of it. The check is excluded from the score rather than counted as a failure.
Fix. Publish `datePublished` and `dateModified` in the page's structured data as ISO 8601 values with a timezone offset. Update `dateModified` only when the content actually changes, since bumping it on every deploy trains crawlers to ignore it. Mirror the value in a visible `<time datetime>` element so the rendered text and the metadata agree, and keep the sitemap `lastmod` consistent with it.
<time datetime="2026-08-04T14:20:00-04:00">Updated August 4, 2026</time>
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "Allow AI answer engines in robots.txt",
"datePublished": "2026-02-11T09:00:00-05:00",
"dateModified": "2026-08-04T14:20:00-04:00"
}
</script>ReferenceEvidence raw measurements
- Size
- 7,464 bytes
- Groups
- 5
- Sitemaps
- https://www.yhprn.com/sitemap.xml
View the file as our crawler received it
Host: www.yhprn.com
Sitemap: https://www.yhprn.com/sitemap.xml
User-agent: YandexBot
Crawl-delay: 1
Disallow: /cgi-bin/
Disallow: /view/
Disallow: /2257
Disallow: /contact/
Disallow: /embed/
Disallow: /v/
User-agent: bingbot
Crawl-delay: 1
Disallow: /cgi-bin/
Disallow: /view/
Disallow: /2257
Disallow: /contact/
Disallow: /embed/
Disallow: /v/
User-agent: msnbot
Crawl-delay: 1
Disallow: /cgi-bin/
Disallow: /view/
Disallow: /2257
Disallow: /contact/
Disallow: /embed/
Disallow: /v/
User-agent: AbachoBOT
User-agent: AhrefsBot
User-agent: anarchie
User-agent: antibot
User-agent: appie
User-agent: ASPSeek
User-agent: asterias
User-agent: attach
User-agent: autoemailspider
User-agent: B2w
User-agent: BackDoorBot
User-agent: BackWeb
User-agent: Bandit
User-agent: BatchFTP
User-agent: Black\ Hole
User-agent: BlackWidow
User-agent: BlowFish
User-agent: Bot\ mailto
User-agent: Bot\ mailto:craftbot@yahoo.com
User-agent: BotALot
User-agent: Buddy
User-agent: BuiltBotTough
User-agent: Bullseye
User-agent: bumblebee
User-agent: BunnySlippers
User-agent: CheeseBot
User-agent: CherryPicker
User-agent: CherryPickerElite
User-agent: CherryPickerSE
User-agent: ChinaClaw
User-agent: ClariaBot
User-agent: clsHTTP
User-agent: COAST\ WebMaster
User-agent: ColdFusion
User-agent: Collector
User-agent: Copier
User-agent: CopyRightCheck
User-agent: cosmos
User-agent: Crescent
User-agent: curl
User-agent: Custo
User-agent: DA
User-agent: Diamond
User-agent: DISCo
User-agent: DISCo\ Pump
User-agent: DittoSpyder
User-agent: dloader
User-agent: Dotbot
User-agent: Download\ Demon
User-agent: Download\ Wonder
User-agent: Downloader
User-agent: Drip
User-agent: DTS\ Agent
User-agent: EasyDL
User-agent: eCatch
User-agent: EirGrabber
User-agent: EmailCollector
User-agent: EmailSiphon
User-agent: EmailWolf
User-agent: EroCrawler
User-agent: Exabot
User-agent: Express\ WebPictures
User-agent: ExtractorPro
User-agent: Extreme\ Picture\ Finder
User-agent: EyeNetIE
User-agent: FAST\ WebCrawler
User-agent: Fetch\ API\ Request
User-agent: FileHound
User-agent: FlashGet
User-agent: FlickBot
User-agent: FreeFind.com
User-agent: FrontPage
User-agent: Generic
User-agent: GetRight
User-agent: GetSmart
User-agent: GetWeb!
User-agent: Gigabot
User-agent: Go!Zilla
User-agent: Go-Ahead-Got-It
User-agent: gotit
User-agent: Grabber
User-agent: GrabNet
User-agent: Grafula
User-agent: Gulliver
User-agent: Harvest
User-agent: Heretrix
User-agent: HitboxDoctor
User-agent: hloader
User-agent: HMView
User-agent: HTTPapp
User-agent: httpfetcher
User-agent: httplib
User-agent: httpscraper
User-agent: HTTPTrack
User-agent: HTTPviewer
User-agent: HTTrack
User-agent: humanlinks
User-agent: ia_archiver
User-agent: Image\ Stripper
User-agent: 360Spider
User-agent: CuteStat
User-agent: Image\ Sucker
User-agent: Indy\ Library
User-agent: InfoNaviRobot
User-agent: InterGET
User-agent: Internet\ Ninja
User-agent: InternetSeer.com
User-agent: Iria
User-agent: IRLbot
User-agent: Java
User-agent: JennyBot
User-agent: JetCar
User-agent: JoBo
User-agent: JOC
User-agent: JOC\ Web\ Spider
User-agent: Jonzilla
User-agent: JustView
User-agent: Kenjin\ Spider
User-agent: Keyword\ Density
User-agent: Lachesis
User-agent: larbin
User-agent: LeechFTP
User-agent: LexiBot
User-agent: lftp
User-agent: Libby_
User-agent: libWeb
User-agent: libwww-perl
User-agent: libwwwperl
User-agent: likse
User-agent: Link
User-agent: LinkextractorPro
User-agent: LinkScan
User-agent: LinkWalker
User-agent: lwp-trivial
User-agent: lwp\ request
User-agent: Mag-Net
User-agent: Magnet
User-agent: Mass\ Downloader
User-agent: Mata\ Hari
User-agent: Memo
User-agent: Mercator
User-agent: Metacarta
User-agent: Mewsoft\ Search\ Engine
User-agent: MFC_Tear_Sample
User-agent: Microsoft\ URL\ Control
User-agent: MicrosoftURL
User-agent: MIDown\ tool
User-agent: MIIxpc
User-agent: Mirror
User-agent: Missigua
User-agent: Mister\ PiX
User-agent: MJ12bot
User-agent: moget
User-agent: MSFrontPage
User-agent: MSIECrawler
User-agent: NationalDirectory\ WebSpider
User-agent: Navroad
User-agent: NearSite
User-agent: Net\ Probe
User-agent: Net\ Vampire
User-agent: NetAnts
User-agent: NetMechanic
User-agent: NetResearchServer
User-agent: NetSpider
User-agent: NetZip
User-agent: NetZIP
User-agent: nexuscache
User-agent: NICErsPRO
User-agent: Nikto
User-agent: Ninja
User-agent: NPBot
User-agent: oBot
User-agent: Octopus
User-agent: Offline\ Explorer
User-agent: Offline\ Navigator
User-agent: onestop
User-agent: Openfind
User-agent: Openfind\ data\ gatherer
User-agent: OrangeBot
User-agent: our\ agent
User-agent: PageGrabber
User-agent: Papa\ Foto
User-agent: pavuk
User-agent: pcBrowser
User-agent: Perl
User-agent: PHP
User-agent: PHP\ version
User-agent: PHPot
User-agent: Ping
User-agent: PingALink\ Monitoring\ Services
User-agent: Pockey
User-agent: Pompos
User-agent: ProPowerBot
User-agent: ProWebWalker
User-agent: psbot
User-agent: psycheclone
User-agent: Pump
User-agent: Python-urllib
User-agent: BUbiNG
User-agent: ltx71 - (http://ltx71.com/)
User-agent: Python\ urllib
User-agent: QueryN
User-agent: RealDownload
User-agent: Reaper
User-agent: Recorder
User-agent: ReGet
User-agent: RepoMonkey
User-agent: Rico
User-agent: RMA
User-agent: Robozilla
User-agent: Rogerbot
User-agent: Scooter
User-agent: ScoutAbout
User-agent: SemrushBot-SA
User-agent: Siphon
User-agent: sitecheck.internetseer.com
User-agent: SiteSnagger
User-agent: slysearch
User-agent: SmartDownload
User-agent: Snake
User-agent: Snapbot
User-agent: Snoopy
User-agent: SpaceBison
User-agent: SpankBot
User-agent: spanner
User-agent: Spinne
User-agent: Sqworm
User-agent: Stealer
User-agent: Stripper
User-agent: Sucker
User-agent: SuperBot
User-agent: SuperHTTP
User-agent: Surfbot
User-agent: suzuran
User-agent: Szukacz
User-agent: tAkeOut
User-agent: Teleport\ Pro
User-agent: Telesoft
User-agent: The\ Intraformant
User-agent: TheNomad
User-agent: TightTwatBot
User-agent: Titan
User-agent: toCrawl
User-agent: True_Robot
User-agent: turingos
User-ag- llms.txt
- absent
- ai.txt
- absent
- Sitemap
- /sitemap.xml (5 URLs)
- Feeds
- none
- Schema types
- none
- Edge
- not identified
- Final URL
- https://yhprn.com/
- HTML size
- 113 KB, 781 words of text
Compare it with a competitor head to head
Put yhprn.com next to another domain and see which one an answer engine can actually read. Both sides are measured the same way, so the difference is the finding.
Already measured: vs bepress.com · vs quip.com · vs fairmont.com
Publish the score free
Sites that score well embed the badge. It links back to this live report, which re-measures on every scan.
<a href="https://crawlcensus.com/site/yhprn.com"><img src="https://crawlcensus.com/badge/yhprn.com.svg" alt="AI access score for yhprn.com" width="150" height="20"></a>Track changes on this domain free for one
Crawler policy is edited quietly. We re-scan monitored domains daily, keep the history, and email you the moment a crawler is blocked or unblocked, an llms.txt appears, or the score moves.
Free for one domain per address, and always free for domains you have verified you own. No newsletter, only change alerts.
Pro follows 25 domains, keeps 90 days of history, compares them against each other and exports the lot as CSV, for 19 dollars a month. That is the difference: measuring is free, watching a portfolio is not.