Grade B 83/10083Grade B
AI access report

bamboohr.com

Open to answer engines, refuses AI user agents at the edge, publishes llms.txt.

4 failing 3 partial 20 passing Scanned 9 hours ago · 2 scans on record
JSON Monitor this site
Reach
34 / 40

Whether AI crawlers and answer-engine fetchers are permitted to request the page at all, in robots.txt, in robots directives, and at the edge.

Readability
22.5 / 25

Whether a fetcher that does not execute JavaScript receives the actual content, in markup an extractor can segment.

Structure
18 / 20

Machine-readable markup that states the page's type, entities, canonical URL, and discrete facts instead of leaving them to be inferred.

Attribution
8 / 15

Signals that let an answer engine name the author, date the content, resolve the publisher, and cite it under known terms.

Who is allowed to read this site 0 of 24 answer engines blocked

CrawlerOperatorUses content forrobots.txtLive request
GPTBot
Crawls content that may be used to train OpenAI's generative AI foundation models.
OpenAI Model training allowed
Allow: /
refused 403
OAI-SearchBot
Indexes pages so they can be surfaced and cited in ChatGPT search results, not for training.
OpenAI Answer index allowed
Allow: /
refused 403
ChatGPT-User
Fetches a page when a ChatGPT user or GPT Action asks for it; user-initiated, so robots rules may not apply.
OpenAI Live retrieval allowed
Allow: /
not probed
OAI-AdsBot
Visits pages submitted as ChatGPT ads to check policy compliance and ad relevance; not used for model training.
OpenAI Live retrieval allowed
Allow: /
not probed
ClaudeBot
Collects web content that may contribute to training Anthropic's models; honors Crawl-delay.
Anthropic Model training allowed
Allow: /
refused 403
Claude-User
Retrieves pages on demand when a Claude user's question needs live web content.
Anthropic Live retrieval allowed
Allow: /
not probed
Claude-SearchBot
Indexes content to improve the relevance and accuracy of Claude's search results.
Anthropic Answer index allowed
Allow: /
not probed
anthropic-ai
Legacy token widely blocked for Anthropic training; Anthropic now documents ClaudeBot, Claude-User and Claude-SearchBot.
Anthropic Model training allowed
Allow: /
not probed
Google-Extended
Control token with no user agent of its own; governs Gemini training and grounding use of Googlebot data.
Google Model training allowed
Allow: /
not probed
Googlebot
Crawls and renders pages for Google Search, Images, Video, News and Discover.
Google Answer index allowed
Allow: /
not probed
Googlebot-News
Robots token controlling Google News inclusion; crawling itself uses the Googlebot user agents.
Google Answer index allowed
Allow: /
not probed
Google-CloudVertexBot
Crawls sites at a site owner's request to build Vertex AI agents; no effect on Google Search.
Google Live retrieval allowed
Allow: /
not probed
GoogleOther
Generic Google crawler used by product teams for one-off fetches such as internal research and development.
Google Model training allowed
Allow: /
not probed
Applebot
Crawls for Siri, Spotlight and Safari search; falls back to Googlebot rules and ignores Crawl-delay.
Apple Answer index allowed
Allow: /
not probed
Applebot-Extended
Control token with no user agent; disallowing it excludes crawled content from Apple foundation model training.
Apple Model training allowed
Allow: /
not probed
Bingbot
Indexes pages for Bing search and the Copilot answers that are grounded in the Bing index.
Microsoft Answer index allowed
Allow: /
not probed
msnbot
Legacy Microsoft search crawler token still honored alongside bingbot.
Microsoft Answer index allowed
Allow: /
not probed
PerplexityBot
Indexes and links pages in Perplexity search results; not used to collect foundation model training data.
Perplexity Answer index allowed
Allow: /
refused 403
Perplexity-User ignores robots
Fetches a page for a specific user question; Perplexity documents that it generally ignores robots.txt.
Perplexity Live retrieval allowed
Allow: /
not probed
Meta-ExternalAgent
Crawls the web to train Meta's foundation AI models and to index content directly into products.
Meta Model training allowed
Allow: /
not probed
Meta-ExternalFetcher ignores robots
Fetches individual links for agentic AI tasks; Meta documents that it may bypass robots.txt.
Meta Live retrieval allowed
Allow: /
not probed
FacebookBot
Crawls public pages to improve language models behind Meta's speech recognition technology.
Meta Model training allowed
Allow: /
not probed
Meta-WebIndexer
Indexes pages so Meta AI can cite and link them in its search answers.
Meta Answer index allowed
Allow: /
not probed
Meta-ExternalAds
Crawls the web to improve Meta's advertising and other business products and services.
Meta Model training allowed
Allow: /
not probed
facebookexternalhit ignores robots
Fetches shared links for Facebook, Instagram and Messenger previews; may bypass robots.txt for integrity checks.
Meta Live retrieval allowed
Allow: /
not probed
Bytespider ignores robots
Downloads content to train ByteDance LLMs and is widely reported to ignore robots.txt directives.
ByteDance Model training allowed
Allow: /
not probed
TikTokSpider ignores robots
Fetches shared URLs for TikTok link previews and feeds; not expected to follow robots.txt.
ByteDance Live retrieval allowed
Allow: /
not probed
Amazonbot
Crawls for Amazon product and Alexa answers and may use the content to train Amazon AI models.
Amazon Model training allowed
Allow: /
not probed
Amzn-SearchBot
Indexes content for Amazon search experiences such as Alexa; does not crawl for generative AI training.
Amazon Answer index allowed
Allow: /
not probed
Amzn-User ignores robots
Fetches live pages to answer a user's Alexa question; Amazon documents it may not follow all robots.txt rules.
Amazon Live retrieval allowed
Allow: /
not probed
CCBot
Builds the open Common Crawl web archive, a common source of LLM pretraining corpora.
Common Crawl Foundation Archive allowed
Allow: /
not probed
Diffbot
Extracts structured page data for Diffbot's knowledge graph, which is licensed to AI customers.
Diffbot Model training allowed
Allow: /
not probed
omgili
Collects forum, news and blog content that Webz.io sells as web data feeds, including for AI training.
Webz.io Model training allowed
Allow: /
not probed
omgilibot
Legacy Omgili search crawler token still blocked alongside the current omgili agent.
Webz.io Model training allowed
Allow: /
not probed
AI2Bot
Collects web text for Ai2's open datasets used to train open language models such as OLMo.
Allen Institute for AI Model training allowed
Allow: /
not probed
cohere-ai
Retrieves pages to answer user-initiated prompts in Cohere's enterprise AI products.
Cohere Live retrieval allowed
Allow: /
not probed
cohere-training-data-crawler
Downloads training data for the large language models behind Cohere's enterprise AI products.
Cohere Model training allowed
Allow: /
not probed
MistralAI-User
Fetches pages on demand so Mistral's Vibe assistant can answer a question with live, cited web content.
Mistral AI Live retrieval allowed
Allow: /
not probed
MistralAI-Index
Indexes content for Mistral search behind Vibe answers; not used for generative AI training.
Mistral AI Answer index allowed
Allow: /
not probed
MistralAI-Training
Crawls web content to build datasets for training Mistral's generative AI models.
Mistral AI Model training allowed
Allow: /
not probed
DuckAssistBot
Crawls pages in real time for DuckDuckGo's cited AI-assisted answers; not used for model training.
DuckDuckGo Live retrieval allowed
Allow: /
not probed
YouBot
Indexes pages for You.com search results and the AI answers built on that index.
You.com Answer index allowed
Allow: /
not probed
PanguBot
Collects web content used to train Huawei's PanGu family of large models.
Huawei Model training allowed
Allow: /
not probed
Timpibot
Crawls pages for Timpi's decentralized index, which is also used as LLM training data.
Timpi Model training allowed
Allow: /
not probed
ImagesiftBot
Downloads public images plus surrounding text to build ImageSift's searchable image index.
ImageSift (Hive) Model training allowed
Allow: /
not probed
Kangaroo Bot
Scrapes site content into datasets used to train the Kangaroo LLM.
Kangaroo LLM Model training allowed
Allow: /
not probed
SemrushBot-OCOB
Crawls pages to feed Semrush's ContentShake AI writing tool.
Semrush Model training allowed
Allow: /
not probed
Scrapy
Generic scraping framework often used to build AI training datasets; obeys robots.txt only when ROBOTSTXT_OBEY is on.
Zyte (open-source framework) Model training allowed
Allow: /
not probed
Two of these columns matter differently. robots.txt is what the site declares. Live request is what actually happened when we sent a real request using that crawler's user agent from a datacentre IP, which is how edge blocking, rate limits and challenge pages show up even when robots.txt looks permissive.

Reach 34 / 40

Edge returns 403, 429, or a challenge to AI user agents

Requests identifying as gptbot, oai-searchbot, perplexitybot, claudebot were refused at the edge (HTTP 403 Forbidden; HTTP 403 Forbidden; HTTP 403 Forbidden; HTTP 403 Forbidden).

Why it matters. A permissive robots.txt is irrelevant if the edge answers an AI user agent with 403, 429, or a JavaScript challenge. Retrieval fetchers do not solve interstitials, so the assistant records a fetch failure and answers from another source.
Fix. Fetch the page with each AI user agent string and compare the status and byte count against a browser request. On Cloudflare, check whether the "Block AI bots" toggle in AI Crawl Control, a Bot Fight Mode rule, or a WAF custom rule on `cf.verified_bot_category` is catching the request, then narrow it: block the training category and add a skip rule for the retrieval agents you want citing you. Verified bots must not be handed Managed Challenge, since a challenge is a hard failure for a non-browser client. Re-test after every WAF or bot-management change, because these toggles are applied zone-wide.
for ua in "OAI-SearchBot/1.0" "ChatGPT-User/1.0" "Claude-User/1.0" \
          "Claude-SearchBot/1.0" "PerplexityBot/1.0" "Mozilla/5.0"; do
  code=$(curl -s -o /dev/null -w '%{http_code}' -A "$ua" https://example.com/)
  printf '%s\t%s\n' "$code" "$ua"
done
Reference
6 pt

Answer-engine fetchers are allowed to retrieve and cite this page

All 24 answer-engine fetchers are allowed to retrieve pages for citation.

12 pt

Training crawlers may fetch this path

All 23 tracked training crawlers are allowed.

8 pt

No blanket disallow applies to this path

The wildcard group does not disallow the entire site.

5 pt

robots.txt served as plain text with a 200 response

robots.txt served, 2891 bytes, 11 group(s).

3 pt

No Crawl-delay directive constrains fetchers

No Crawl-delay directive.

2 pt

Page is indexable, with no noindex directive

No noindex directive on the homepage.

2 pt

Full-length snippet extraction is permitted

Snippets are not restricted by meta tags.

1 pt

X-Robots-Tag header is absent or permissive

No restrictive X-Robots-Tag header.

1 pt

Readability 22.5 / 25

No lang attribute declares the document language

No lang attribute on <html>.

Why it matters. `<html lang>` tells tokenizers and language filters which language the text is in. Without it a retrieval pipeline guesses, and a wrong guess drops the page from language-scoped candidate sets and degrades sentence splitting.
Fix. Set a valid BCP 47 tag on the root element, such as `lang="en"` or `lang="pt-BR"`. Mark inline passages in another language with `lang` on the containing element. If you publish translations, pair the declaration with `hreflang` alternates so each version is attributed to the right locale.
<html lang="en">
  <head>
    <link rel="alternate" hreflang="es" href="https://example.com/es/page" />
  </head>
</html>
Reference
1 pt
!

Title is missing, duplicated, or too generic to identify the page

Title is 70 characters: "BambooHR: The Complete HR Software for People, Payroll &#x26; Benefits"

Why it matters. The `<title>` is the first text an indexer stores for the URL and is frequently reused verbatim as the link label in a generated answer. An empty, duplicated, or purely brand-name title gives the ranker no lexical signal for the page's actual topic.
Fix. Write a unique `<title>` of roughly 30 to 60 characters that leads with the page's subject and ends with the brand. Match the wording of the `h1` so the stored label and the visible heading agree. Remove template output such as "Home" or a bare site name, and keep keyword lists and separator chains out of it.
<title>Allow AI answer engines in robots.txt | Example</title>
Reference
3 pt

Substantive text is present in the server-rendered HTML

958 words of text are present in the raw HTML. Most AI fetchers do not run JavaScript.

9 pt

Visible text makes up a healthy share of the HTML payload

Text is 14.5% of the 43 KB document; 0 KB is inline script.

4 pt

Primary content is wrapped in a semantic landmark

A <main> landmark marks the primary content.

3 pt

Headings form a single, ordered outline

1 H1 and 25 headings total, 7 phrased as questions.

3 pt

Meta description provides an author-written summary

Meta description is 138 characters.

2 pt

Structure 18 / 20

!

Structured data is too generic for what the page is about

Recognized types: organization.

Why it matters. Answer engines route by entity type: a `Product` node supplies price and availability, an `Article` node supplies author and dates, and a `FAQPage` node supplies question and answer pairs. A generic `WebPage` or `WebSite` node on a product or article page carries none of those fields, so the specific facts stay unavailable.
Fix. Replace bare `WebPage` and `WebSite` nodes with the most specific type that describes the page, and fill the properties that type defines. Use `@graph` to publish several linked nodes on one page, such as an `Article` whose `publisher` points at an `Organization` node by `@id`. Add `BreadcrumbList` for hierarchy and reuse the same `@id` values across pages so the entity resolves to one record.
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "Product",
      "@id": "https://example.com/products/widget#product",
      "name": "Widget Pro",
      "sku": "WGT-PRO-1",
      "brand": { "@type": "Brand", "name": "Example" },
      "offers": {
        "@type": "Offer",
        "url": "https://example.com/products/widget",
        "price": "49.00",
        "priceCurrency": "USD",
        "availability": "https://schema.org/InStock"
      }
    },
    {
      "@type": "BreadcrumbList",
      "itemListElement": [
        { "@type": "ListItem", "position": 1, "name": "Products", "item": "https://example.com/products" },
        { "@type": "ListItem", "position": 2, "name": "Widget Pro" }
      ]
    }
  ]
}
Reference
4 pt

Page ships JSON-LD structured data

4 JSON-LD node(s): Organization, ContactPoint, PostalAddress, AggregateRating.

7 pt

JSON-LD parses cleanly with recognised schema.org terms

All JSON-LD blocks parse cleanly.

3 pt

Page declares a self-referential canonical URL

Canonical: https://www.bamboohr.com/

3 pt

XML sitemap is declared in robots.txt and returns 200

Sitemap found at /sitemap.xml (67 URLs on the first document).

2 pt

Key facts are available in lists or tables

0 tables, 1 lists, 0 code blocks, 7 question headings.

1 pt

Attribution 8 / 15

No machine-readable author is attached to the page

No author or Person entity, which weakens the authority signals answer engines use.

Why it matters. An `author` property in structured data is what lets an answer engine name a person or organisation as the source and link the byline to a stable profile. A byline that exists only as styled text is not reliably associated with the document during extraction.
Fix. Add an `author` property to the page's `Article`, `BlogPosting`, or `NewsArticle` node, typed as `Person` or `Organization`, with a `name` and a `url` pointing at a real profile page. Give each author a stable `@id` and reuse it across posts so the entity consolidates. Keep the visible byline identical to the structured value, and avoid generic names such as "Admin" or "Staff Writer" where a real author exists.
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "headline": "Allow AI answer engines in robots.txt",
  "author": {
    "@type": "Person",
    "@id": "https://example.com/authors/dana-reyes#person",
    "name": "Dana Reyes",
    "url": "https://example.com/authors/dana-reyes",
    "jobTitle": "Infrastructure Engineer",
    "sameAs": ["https://github.com/danareyes"]
  }
}
Reference
3 pt

No machine-readable published or modified date

No publication or modification dates in structured data.

Why it matters. Answer engines prefer recent sources for questions about current state and use `dateModified` to decide whether a cached copy needs refetching. With no machine-readable date the page is treated as undated and loses to competitors that publish one.
Fix. Publish `datePublished` and `dateModified` in the page's structured data as ISO 8601 values with a timezone offset. Update `dateModified` only when the content actually changes, since bumping it on every deploy trains crawlers to ignore it. Mirror the value in a visible `<time datetime>` element so the rendered text and the metadata agree, and keep the sitemap `lastmod` consistent with it.
<time datetime="2026-08-04T14:20:00-04:00">Updated August 4, 2026</time>

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Allow AI answer engines in robots.txt",
  "datePublished": "2026-02-11T09:00:00-05:00",
  "dateModified": "2026-08-04T14:20:00-04:00"
}
</script>
Reference
3 pt
!

No machine-readable license or usage terms for the content

No licence declaration, so reuse terms are ambiguous.

Why it matters. A `license` property, or a linked terms page, states the reuse conditions in a place a crawler can read, rather than leaving them to be inferred. Where terms are unstated, some pipelines default to the more restrictive handling, which reduces how much of the text is quoted.
Fix. Add a `license` property to the page's structured data pointing at a specific license URL, such as a Creative Commons deed or your own terms page, and add `rel="license"` on the visible link. Use `usageInfo` for conditions that are not a standard license, such as attribution wording or an API-only clause. State the terms once, at a stable URL, and reference it from every page rather than restating it per template.
<a rel="license" href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a>

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Allow AI answer engines in robots.txt",
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "usageInfo": "https://example.com/legal/content-reuse",
  "creditText": "Example, crawlcensus research desk"
}
</script>
Reference
2 pt

/llms.txt lists canonical pages in the documented format

/llms.txt is present and well formed (5 sections, 25 links).

4 pt

Organization entity declares the publisher and its identifiers

Organization entity with sameAs links is present.

3 pt

Evidence raw measurements

robots.txt
Size
2,891 bytes
Groups
11
Sitemaps
https://www.bamboohr.com/sitemap-index.xml
View the file as our crawler received it
User-agent: *
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /pricing/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/

User-agent: GPTBot
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

User-agent: ClaudeBot
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

User-agent: Google-Extended
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

User-agent: Anthropic-AI
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

User-agent: Bytespider
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

User-agent: CCBot
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

User-agent: Applebot-Extended
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

User-agent: Meta-ExternalAgent
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

User-agent: OAI-SearchBot
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /pricing/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

User-agent: PerplexityBot
Allow: /
Allow: /platform/
Allow: /blog/
Allow: /resources/
Allow: /pricing/
Allow: /why-bamboohr/
Allow: /integrations/
Allow: /customers/
Disallow: /pl-pages/
Disallow: /pl/
Disallow: /lp/
Disallow: /*.plain.html$
Disallow: /blog/fixtures/

Sitemap: https://www.bamboohr.com/sitemap-index.xml
llms: https://www.bamboohr.com/llms.txt
Machine-readable extras
llms.txt
valid, 6,301 bytes, 25 links
ai.txt
absent
Sitemap
/sitemap.xml (67 URLs)
Feeds
none
Schema types
Organization, ContactPoint, PostalAddress, AggregateRating
Edge
Cloudflare
Final URL
https://www.bamboohr.com/
HTML size
43 KB, 958 words of text

Publish the score free

Sites that score well embed the badge. It links back to this live report, which re-measures on every scan.

AI access score badge
<a href="https://crawlcensus.com/site/bamboohr.com"><img src="https://crawlcensus.com/badge/bamboohr.com.svg" alt="AI access score for bamboohr.com" width="150" height="20"></a>

Track changes on this domain pro

Crawler policy is edited quietly. We re-scan monitored domains daily, keep the history, and email you the moment a crawler is blocked or unblocked, an llms.txt appears, or the score moves.

One address, unlimited domains during the beta. No newsletter, only change alerts.