Grade B 74/10074Grade B
AI access report

courrierinternational.com

Blocks 7 of 24 answer engines, blocks 19 training crawlers.

5 failing 5 partial 17 passing Scanned 5 hours ago · 1 scan on record
JSON Claim this domain
Reach
28.9 / 40

Whether AI crawlers and answer-engine fetchers are permitted to request the page at all. What happens on the wire outweighs what the policy file says: a live refusal costs 16 of these 40 points, twice what the heaviest robots.txt rule costs, because a page a crawler cannot fetch is unreadable no matter how well it is marked up.

Readability
25 / 25

Whether a fetcher that does not execute JavaScript receives the actual content, in markup an extractor can segment.

Structure
15.5 / 20

Machine-readable markup that states the page's type, entities, canonical URL, and discrete facts instead of leaving them to be inferred.

Attribution
4 / 15

Signals that let an answer engine name the author, date the content, resolve the publisher, and cite it under known terms.

Who is allowed to read this site 7 of 24 answer engines blocked

CrawlerOperatorUses content forrobots.txtLive request
GPTBot
Crawls content that may be used to train OpenAI's generative AI foundation models.
OpenAI Model training blocked
Disallow: /
served 200
OAI-SearchBot
Indexes pages so they can be surfaced and cited in ChatGPT search results, not for training.
OpenAI Answer index allowed served 200
ChatGPT-User
Fetches a page when a ChatGPT user or GPT Action asks for it; user-initiated, so robots rules may not apply.
OpenAI Live retrieval blocked
Disallow: /
not probed
OAI-AdsBot
Visits pages submitted as ChatGPT ads to check policy compliance and ad relevance; not used for model training.
OpenAI Live retrieval allowed not probed
ClaudeBot
Collects web content that may contribute to training Anthropic's models; honors Crawl-delay.
Anthropic Model training blocked
Disallow: /
served 200
Claude-User
Retrieves pages on demand when a Claude user's question needs live web content.
Anthropic Live retrieval allowed not probed
Claude-SearchBot
Indexes content to improve the relevance and accuracy of Claude's search results.
Anthropic Answer index allowed not probed
anthropic-ai
Legacy token widely blocked for Anthropic training; Anthropic now documents ClaudeBot, Claude-User and Claude-SearchBot.
Anthropic Model training blocked
Disallow: /
not probed
Google-Extended
Control token with no user agent of its own; governs Gemini training and grounding use of Googlebot data.
Google Model training blocked
Disallow: /
not probed
Googlebot
Crawls and renders pages for Google Search, Images, Video, News and Discover.
Google Answer index allowed not probed
Googlebot-News
Robots token controlling Google News inclusion; crawling itself uses the Googlebot user agents.
Google Answer index allowed not probed
Google-CloudVertexBot
Crawls sites at a site owner's request to build Vertex AI agents; no effect on Google Search.
Google Live retrieval allowed not probed
GoogleOther
Generic Google crawler used by product teams for one-off fetches such as internal research and development.
Google Model training allowed not probed
Applebot
Crawls for Siri, Spotlight and Safari search; falls back to Googlebot rules and ignores Crawl-delay.
Apple Answer index blocked
Disallow: /
not probed
Applebot-Extended
Control token with no user agent; disallowing it excludes crawled content from Apple foundation model training.
Apple Model training blocked
Disallow: /
not probed
Bingbot
Indexes pages for Bing search and the Copilot answers that are grounded in the Bing index.
Microsoft Answer index allowed not probed
msnbot
Legacy Microsoft search crawler token still honored alongside bingbot.
Microsoft Answer index allowed not probed
PerplexityBot
Indexes and links pages in Perplexity search results; not used to collect foundation model training data.
Perplexity Answer index blocked
Disallow: /
served 200
Perplexity-User ignores robots
Fetches a page for a specific user question; Perplexity documents that it generally ignores robots.txt.
Perplexity Live retrieval allowed not probed
Meta-ExternalAgent
Crawls the web to train Meta's foundation AI models and to index content directly into products.
Meta Model training blocked
Disallow: /
not probed
Meta-ExternalFetcher ignores robots
Fetches individual links for agentic AI tasks; Meta documents that it may bypass robots.txt.
Meta Live retrieval blocked
Disallow: /
not probed
FacebookBot
Crawls public pages to improve language models behind Meta's speech recognition technology.
Meta Model training blocked
Disallow: /
not probed
Meta-WebIndexer
Indexes pages so Meta AI can cite and link them in its search answers.
Meta Answer index allowed not probed
Meta-ExternalAds
Crawls the web to improve Meta's advertising and other business products and services.
Meta Model training allowed not probed
facebookexternalhit ignores robots
Fetches shared links for Facebook, Instagram and Messenger previews; may bypass robots.txt for integrity checks.
Meta Live retrieval allowed not probed
Bytespider ignores robots
Downloads content to train ByteDance LLMs and is widely reported to ignore robots.txt directives.
ByteDance Model training blocked
Disallow: /
not probed
TikTokSpider ignores robots
Fetches shared URLs for TikTok link previews and feeds; not expected to follow robots.txt.
ByteDance Live retrieval allowed not probed
Amazonbot
Crawls for Amazon product and Alexa answers and may use the content to train Amazon AI models.
Amazon Model training blocked
Disallow: /
not probed
Amzn-SearchBot
Indexes content for Amazon search experiences such as Alexa; does not crawl for generative AI training.
Amazon Answer index allowed not probed
Amzn-User ignores robots
Fetches live pages to answer a user's Alexa question; Amazon documents it may not follow all robots.txt rules.
Amazon Live retrieval allowed not probed
CCBot
Builds the open Common Crawl web archive, a common source of LLM pretraining corpora.
Common Crawl Foundation Archive blocked
Disallow: /
not probed
Diffbot
Extracts structured page data for Diffbot's knowledge graph, which is licensed to AI customers.
Diffbot Model training blocked
Disallow: /
not probed
omgili
Collects forum, news and blog content that Webz.io sells as web data feeds, including for AI training.
Webz.io Model training blocked
Disallow: /
not probed
omgilibot
Legacy Omgili search crawler token still blocked alongside the current omgili agent.
Webz.io Model training blocked
Disallow: /
not probed
AI2Bot
Collects web text for Ai2's open datasets used to train open language models such as OLMo.
Allen Institute for AI Model training blocked
Disallow: /
not probed
cohere-ai
Retrieves pages to answer user-initiated prompts in Cohere's enterprise AI products.
Cohere Live retrieval blocked
Disallow: /
not probed
cohere-training-data-crawler
Downloads training data for the large language models behind Cohere's enterprise AI products.
Cohere Model training blocked
Disallow: /
not probed
MistralAI-User
Fetches pages on demand so Mistral's Vibe assistant can answer a question with live, cited web content.
Mistral AI Live retrieval allowed not probed
MistralAI-Index
Indexes content for Mistral search behind Vibe answers; not used for generative AI training.
Mistral AI Answer index allowed not probed
MistralAI-Training
Crawls web content to build datasets for training Mistral's generative AI models.
Mistral AI Model training allowed not probed
DuckAssistBot
Crawls pages in real time for DuckDuckGo's cited AI-assisted answers; not used for model training.
DuckDuckGo Live retrieval blocked
Disallow: /
not probed
YouBot
Indexes pages for You.com search results and the AI answers built on that index.
You.com Answer index blocked
Disallow: /
not probed
PanguBot
Collects web content used to train Huawei's PanGu family of large models.
Huawei Model training blocked
Disallow: /
not probed
Timpibot
Crawls pages for Timpi's decentralized index, which is also used as LLM training data.
Timpi Model training blocked
Disallow: /
not probed
ImagesiftBot
Downloads public images plus surrounding text to build ImageSift's searchable image index.
ImageSift (Hive) Model training blocked
Disallow: /
not probed
Kangaroo Bot
Scrapes site content into datasets used to train the Kangaroo LLM.
Kangaroo LLM Model training blocked
Disallow: /
not probed
SemrushBot-OCOB
Crawls pages to feed Semrush's ContentShake AI writing tool.
Semrush Model training blocked
Disallow: /
not probed
Scrapy
Generic scraping framework often used to build AI training datasets; obeys robots.txt only when ROBOTSTXT_OBEY is on.
Zyte (open-source framework) Model training allowed not probed
Two of these columns matter differently. robots.txt is what the site declares. Live request is what actually happened when we sent a real request using that crawler's user agent from a datacentre IP, which is how edge blocking, rate limits and challenge pages show up even when robots.txt looks permissive.

Reach 28.9 / 40

Answer-engine fetchers are blocked, so you cannot be cited

Blocked from citing you: ChatGPT-User, Applebot, PerplexityBot, Meta-ExternalFetcher, cohere-ai, DuckAssistBot, YouBot.

Why it matters. Retrieval fetchers (OAI-SearchBot, ChatGPT-User, Claude-User, Claude-SearchBot, PerplexityBot, Perplexity-User) load a page at answer time to quote and link it. When robots.txt disallows them the assistant drops the URL from its candidate set, so the citation goes to a competitor while your training exposure stays exactly the same.
Fix. Separate the two crawler classes in robots.txt instead of blocking by vendor. Allow `OAI-SearchBot` and `ChatGPT-User` (OpenAI retrieval and user-initiated fetches), `Claude-SearchBot` and `Claude-User` (Anthropic retrieval), and `PerplexityBot` and `Perplexity-User`; keep any opt-out you want on `GPTBot`, `ClaudeBot`, `CCBot`, `Google-Extended`, and `Applebot-Extended`. Order does not decide precedence in RFC 9309 parsers, the longest matching rule does, so make the allow rules at least as specific as the disallow rules. Blocking retrieval buys nothing on training, because the training crawlers are separate user agents with separate rules.
# Retrieval and citation fetchers: allowed.
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

# Training and dataset crawlers: disallowed.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

User-agent: *
Allow: /
Disallow: /checkout

Sitemap: https://example.com/sitemap.xml
Reference
8 pt
!

Training crawlers are disallowed for this path

Blocked by robots.txt: GPTBot, ClaudeBot, anthropic-ai, Google-Extended, Applebot-Extended, Meta-ExternalAgent, FacebookBot, Bytespider, Amazonbot, Diffbot, omgili, omgilibot, AI2Bot, cohere-training-data-crawler, PanguBot, Timpibot, ImagesiftBot, Kangaroo Bot, SemrushBot-OCOB.

Why it matters. GPTBot, ClaudeBot, and CCBot collect pages into pretraining corpora, and Google-Extended and Applebot-Extended are opt-out tokens that govern whether already-crawled pages may be used for Gemini and Apple Intelligence. Disallowing them removes your text from the corpora models generalise from, which is a policy choice, not a bug.
Fix. Decide this deliberately rather than by inheriting a template. If you want your content in model weights, remove the `Disallow: /` groups for `GPTBot`, `ClaudeBot`, `CCBot`, `Google-Extended`, and `Applebot-Extended`. If you do not, keep those groups but scope them to the paths that matter and leave retrieval fetchers untouched, because `Google-Extended` and `Applebot-Extended` only control training use and never affect search or answer citation. Record the decision somewhere durable so the next robots.txt edit does not silently reverse it.
# Opt out of model training only.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
Reference
5 pt
!

robots.txt sets Crawl-delay, which most fetchers ignore

Crawl-delay of 10s is declared. Large crawlers may fetch only a few pages per hour, or ignore the directive entirely.

Why it matters. `Crawl-delay` is not in RFC 9309. Googlebot and OpenAI's fetchers ignore it, while some retrieval agents that do honour it will space requests seconds apart, which can exceed an answer engine's fetch timeout and leave the page uncited.
Fix. Remove `Crawl-delay` and handle load at the edge instead, with caching and rate limiting keyed on the client. If crawl volume is the real problem, cache HTML at your CDN so repeat fetches never reach the origin. Keep the file to standard directives (`User-agent`, `Allow`, `Disallow`, `Sitemap`) so parser behaviour is predictable.
# Remove non-standard throttling directives:
# Crawl-delay: 10

User-agent: *
Allow: /
Reference
1 pt

AI user agents receive the same 200 response as browsers

Live requests as 4 AI user agents were served normally.

16 pt

No blanket disallow applies to this path

The wildcard group does not disallow the entire site.

4 pt

robots.txt served as plain text with a 200 response

robots.txt served, 6985 bytes, 3 group(s).

2 pt

Page is indexable, with no noindex directive

No noindex directive on the homepage.

2 pt

Full-length snippet extraction is permitted

Snippets are not restricted by meta tags.

1 pt

X-Robots-Tag header is absent or permissive

No restrictive X-Robots-Tag header.

1 pt

Readability 25 / 25

Substantive text is present in the server-rendered HTML

2759 words of text are present in the raw HTML. Most AI fetchers do not run JavaScript.

9 pt

Visible text makes up a healthy share of the HTML payload

Text is 8.5% of the 229 KB document; 3 KB is inline script.

4 pt

Primary content is wrapped in a semantic landmark

A <main> landmark marks the primary content.

3 pt

Headings form a single, ordered outline

1 H1 and 15 headings total, 1 phrased as questions.

3 pt

Title is unique and describes the page in specific terms

Title is 51 characters: "Courrier international - Actualités France et Monde"

3 pt

Meta description provides an author-written summary

Meta description is 126 characters.

2 pt

Document language is declared on the html element

Declared language: fr.

1 pt

Structure 15.5 / 20

No reachable XML sitemap is declared

No sitemap.xml and none declared in robots.txt.

Why it matters. A sitemap gives crawlers the URL list and `lastmod` timestamps directly, instead of leaving discovery to link traversal that never reaches pages behind search forms or JavaScript routers. Fetchers use `lastmod` to prioritise recrawls, so fresh content is picked up sooner.
Fix. Publish a sitemap of canonical, indexable URLs and reference it with an absolute `Sitemap:` line in robots.txt. Keep each file under 50,000 URLs and 50 MiB uncompressed, using a sitemap index when you exceed either limit. Set `lastmod` from real content changes rather than the build clock, and exclude redirects, error pages, and non-canonical variants.
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://example.com/blog/robots-for-ai</loc>
    <lastmod>2026-08-04T14:20:00-04:00</lastmod>
  </url>
</urlset>
Reference
2 pt
!

Structured data is too generic for what the page is about

Recognized types: website.

Why it matters. Answer engines route by entity type: a `Product` node supplies price and availability, an `Article` node supplies author and dates, and a `FAQPage` node supplies question and answer pairs. A generic `WebPage` or `WebSite` node on a product or article page carries none of those fields, so the specific facts stay unavailable.
Fix. Replace bare `WebPage` and `WebSite` nodes with the most specific type that describes the page, and fill the properties that type defines. Use `@graph` to publish several linked nodes on one page, such as an `Article` whose `publisher` points at an `Organization` node by `@id`. Add `BreadcrumbList` for hierarchy and reuse the same `@id` values across pages so the entity resolves to one record.
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "Product",
      "@id": "https://example.com/products/widget#product",
      "name": "Widget Pro",
      "sku": "WGT-PRO-1",
      "brand": { "@type": "Brand", "name": "Example" },
      "offers": {
        "@type": "Offer",
        "url": "https://example.com/products/widget",
        "price": "49.00",
        "priceCurrency": "USD",
        "availability": "https://schema.org/InStock"
      }
    },
    {
      "@type": "BreadcrumbList",
      "itemListElement": [
        { "@type": "ListItem", "position": 1, "name": "Products", "item": "https://example.com/products" },
        { "@type": "ListItem", "position": 2, "name": "Widget Pro" }
      ]
    }
  ]
}
Reference
4 pt
!

Facts are only in prose, with no list or table structure

0 tables, 0 lists, 0 code blocks, 1 question headings.

Why it matters. Lists, tables with header cells, and definition lists survive HTML-to-text conversion as discrete rows a model can lift into an answer verbatim. The same facts written as a paragraph, or laid out with positioned `div`s, lose their row and column relationships during extraction.
Fix. Put specifications, comparisons, pricing, and steps into real `<table>` markup with `<caption>` and `<th scope>`, or into `<ul>`, `<ol>`, and `<dl>` elements. Do not simulate tables with `div` grids, and avoid images of tables, which carry no extractable text. Keep one fact per row or list item so a chunk stays meaningful on its own.
<table>
  <caption>AI fetcher purposes</caption>
  <thead>
    <tr><th scope="col">User agent</th><th scope="col">Purpose</th></tr>
  </thead>
  <tbody>
    <tr><td>GPTBot</td><td>Model training</td></tr>
    <tr><td>OAI-SearchBot</td><td>Retrieval and citation</td></tr>
  </tbody>
</table>
Reference
1 pt

Page ships JSON-LD structured data

3 JSON-LD node(s): WebSite, SearchAction, NewsMediaOrganization.

7 pt

JSON-LD parses cleanly with recognised schema.org terms

All JSON-LD blocks parse cleanly.

3 pt

Page declares a self-referential canonical URL

Canonical: https://www.courrierinternational.com

3 pt

Attribution 4 / 15

No /llms.txt index of canonical pages

No /llms.txt.

Why it matters. `/llms.txt` is a markdown file that points an assistant at the canonical pages for a site, so retrieval does not depend on which page a search happened to return. Its format is fixed: one `#` title, a `>` blockquote summary, then `##` sections of markdown links with short notes.
Fix. Publish `/llms.txt` as `text/plain` markdown: an `#` H1 with the project name, a `>` blockquote summary, optional plain paragraphs of context, then `##` sections whose bullets are `[title](absolute-url): note`. Link the pages you want quoted, put lower-priority links under an `## Optional` section, and prefer URLs that also serve clean markdown. Keep it generated from the same source as your sitemap so it does not drift, and remember it is a hint for assistants, not an access control mechanism.
# Example

> Example publishes reference documentation for the Widget API and guides for
> configuring crawler access.

Prefer the pages below over search results; each URL is canonical.

## Docs

- [Widget API reference](https://example.com/docs/api): endpoints, auth, limits.
- [Quickstart](https://example.com/docs/quickstart): first request in five minutes.

## Policies

- [Crawler policy](https://example.com/legal/crawlers): which agents we allow.

## Optional

- [Changelog](https://example.com/changelog): dated release notes.
Reference
4 pt

No machine-readable author is attached to the page

No author or Person entity, which weakens the authority signals answer engines use.

Why it matters. An `author` property in structured data is what lets an answer engine name a person or organisation as the source and link the byline to a stable profile. A byline that exists only as styled text is not reliably associated with the document during extraction.
Fix. Add an `author` property to the page's `Article`, `BlogPosting`, or `NewsArticle` node, typed as `Person` or `Organization`, with a `name` and a `url` pointing at a real profile page. Give each author a stable `@id` and reuse it across posts so the entity consolidates. Keep the visible byline identical to the structured value, and avoid generic names such as "Admin" or "Staff Writer" where a real author exists.
{
  "@context": "https://schema.org",
  "@type": "BlogPosting",
  "headline": "Allow AI answer engines in robots.txt",
  "author": {
    "@type": "Person",
    "@id": "https://example.com/authors/dana-reyes#person",
    "name": "Dana Reyes",
    "url": "https://example.com/authors/dana-reyes",
    "jobTitle": "Infrastructure Engineer",
    "sameAs": ["https://github.com/danareyes"]
  }
}
Reference
3 pt

No machine-readable published or modified date

No publication or modification dates in structured data.

Why it matters. Answer engines prefer recent sources for questions about current state and use `dateModified` to decide whether a cached copy needs refetching. With no machine-readable date the page is treated as undated and loses to competitors that publish one.
Fix. Publish `datePublished` and `dateModified` in the page's structured data as ISO 8601 values with a timezone offset. Update `dateModified` only when the content actually changes, since bumping it on every deploy trains crawlers to ignore it. Mirror the value in a visible `<time datetime>` element so the rendered text and the metadata agree, and keep the sitemap `lastmod` consistent with it.
<time datetime="2026-08-04T14:20:00-04:00">Updated August 4, 2026</time>

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Allow AI answer engines in robots.txt",
  "datePublished": "2026-02-11T09:00:00-05:00",
  "dateModified": "2026-08-04T14:20:00-04:00"
}
</script>
Reference
3 pt
!

No machine-readable license or usage terms for the content

No licence declaration, so reuse terms are ambiguous.

Why it matters. A `license` property, or a linked terms page, states the reuse conditions in a place a crawler can read, rather than leaving them to be inferred. Where terms are unstated, some pipelines default to the more restrictive handling, which reduces how much of the text is quoted.
Fix. Add a `license` property to the page's structured data pointing at a specific license URL, such as a Creative Commons deed or your own terms page, and add `rel="license"` on the visible link. Use `usageInfo` for conditions that are not a standard license, such as attribution wording or an API-only clause. State the terms once, at a stable URL, and reference it from every page rather than restating it per template.
<a rel="license" href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</a>

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Allow AI answer engines in robots.txt",
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "usageInfo": "https://example.com/legal/content-reuse",
  "creditText": "Example, crawlcensus research desk"
}
</script>
Reference
2 pt

Organization entity declares the publisher and its identifiers

Organization entity with sameAs links is present.

3 pt

Evidence raw measurements

robots.txt
Size
6,985 bytes
Groups
3
Sitemaps
none declared
View the file as our crawler received it
#
# robots.txt
#
# This file is to prevent the crawling and indexing of certain parts
# of your site by web crawlers and spiders run by sites like Yahoo!
# and Google. By telling these "robots" where not to go on your site,
# you save bandwidth and server resources.
#
# This file will be ignored unless it is at the root of your host:
# Used:    http://example.com/robots.txt
# Ignored: http://example.com/site/robots.txt
#
# For more information about the robots.txt standard, see:
# http://www.robotstxt.org/robotstxt.html
#
# For syntax checking, see:
# http://www.frobee.com/robots-txt-check

User-agent: *
Crawl-delay: 10
# Directories
Disallow: /includes/
Disallow: /misc/
Disallow: /modules/
Disallow: /profiles/
Disallow: /scripts/
Disallow: /themes/
# Files
Disallow: /CHANGELOG.txt
Disallow: /cron.php
Disallow: /INSTALL.mysql.txt
Disallow: /INSTALL.pgsql.txt
Disallow: /INSTALL.sqlite.txt
Disallow: /install.php
Disallow: /INSTALL.txt
Disallow: /LICENSE.txt
Disallow: /MAINTAINERS.txt
Disallow: /update.php
Disallow: /UPGRADE.txt
Disallow: /xmlrpc.php
# Paths (clean URLs)
Disallow: /admin/
Disallow: /comment/reply/
Disallow: /filter/tips/
Disallow: /node/add/
Disallow: /ci/async/
Disallow: /search/
Disallow: /users/
Disallow: /user/register
Disallow: /user/password
Disallow: /user/login
Disallow: /user/logout
Disallow: /inscription
Disallow: /login?
Disallow: /sfuser/deconnexion
Disallow: /?score*
Disallow: /jeu/puzzle/score*
Disallow: /article/audio-stream*
# Paths (no clean URLs)
Disallow: /?q=admin/
Disallow: /?q=comment/reply/
Disallow: /?q=filter/tips/
Disallow: /?q=node/add/
Disallow: /?q=search/
Disallow: /?q=user/password/
Disallow: /?q=user/register/
Disallow: /?q=user/login/
Disallow: /?q=user/logout/

# Robots exclus de toute indexation.
User-agent: 5erue
User-agent: adequat
User-agent: adequat-systems
User-agent: AI2Bot
User-agent: Amazonbot
User-agent: amazon-kendra
User-agent: amazon-QBusiness
User-agent: anthropic-ai
User-agent: Applebot
User-agent: Applebot-Extended
User-agent: asknread.com
User-agent: Bytespider
User-agent: CCBot
User-agent: ChatGPT-User
User-agent: Cision
User-agent: Claude-Web
User-agent: ClaudeBot
User-agent: coexel
User-agent: cohere-training-data-crawler
User-agent: Diffbot
User-agent: DuckAssistBot
User-agent: FacebookBot
User-agent: flipboard
User-agent: Google-Extended
User-agent: GPTBot
User-agent: grub-client
User-agent: infoseek
User-agent: Jetbot
User-agent: k2spider
User-agent: Kangaroo Bot
User-agent: kbcrawl
User-agent: leadbox
User-agent: libwww
User-agent: linkfluence
User-agent: Meltwater
User-agent: mention
User-agent: Meta-ExternalAgent
User-agent: Meta-ExternalFetcher
User-agent: MSIECrawler
User-agent: mytwip
User-agent: Newzbin
User-agent: Offline Explorer
User-Agent: omgili
User-Agent: omgilibot
User-agent: opinion-tracker
User-agent: PanguBot
User-Agent: PerplexityBot
User-agent: proxem
User-agent: Qwam content intelligence
User-agent: scoop.it
User-agent: score3
User-agent: sitecheck.internetseer.com
User-agent: Synthesio
User-agent: Talkwater
User-agent: Teleport
User-agent: TeleportPro
User-agent: Timpibot
User-agent: trendybuzz
User-agent: vecteurplus
User-agent: verticalsearch
User-agent: vsw
User-agent: WebCopier
User-agent: WebStripper
User-agent: Webzio-Extended
User-agent: wget
User-agent: winello
User-agent: YouBot
User-agent: Youmag
User-agent: Zealbot
Disallow: /

# LISTEROBOTS1802
User-agent: 5emeRue
User-agent: ACQUIRE MEDIA
User-agent: ACTIV Financial (CME Group)
User-agent: AlphaSense
User-agent: AmiSoftware
User-agent: archive.org_bot
User-agent: Archive-It
User-agent: ArgClrInt
User-agent: Ask n read
User-agent: Augure
User-agent: auramundi
User-agent: AwarioRssBot
User-agent: AwarioSmartBot
User-agent: Barchart.com
User-agent: BattleFin
User-agent: Bernin IT
User-agent: Blackboard Safeassign
User-agent: BLP_bbot
User-agent: bluematrix
User-agent: Brandwatch
User-agent: Briefcase.news
User-agent: Buck
User-agent: Bytespider
User-agent: CCBo
User-agent: CikisiBot
User-agent: Coexel
User-agent: cohere-ai
User-agent: Comtex News Network
User-agent: ConveraCrawler
User-agent: Copyright Licensing Agency
User-agent: Corporama
User-agent: D&B Hoovers
User-agent: Data Expression
User-agent: Data Observer
User-agent: Dataminr
User-agent: Dealogic
User-agent: Diffbot
User-agent: Digimind
User-agent: DirectFN
User-agent: Dun & Bradstreet
User-agent: Dun & Bradstreet - D&B ESG Intelligence
User-agent: Dun & Bradstreet Data Marketplace
User-agent: Eagle Alpha
User-agent: ecoresearch
User-agent: ellisphere
User-agent: FactSet
User-agent: FeedCheck
User-agent: FeedReader
User-agent: Feedspot
User-agent: Fitch Solutions
User-agent: Founder Apabi
User-agent: Freshbot
User-agent: FriendlyCrawler
User-agent: Gnowit
User-agent: GnowitNewsbot
User-agent: Ground News
User-agent: ia_archiver
User-agent: ICE Connect Desktop Solution
User-agent: ICE Data Services
User-agent: IHS Markit
User-agent: ImageSift
User-agent: InMédia Technologies
User-agent: Innguma
User-agent: Inoreader
User-agent: ISI Emerging Markets
User-agent: KB Crawl SAS
User-agent: Knowings
User-agent: Koyfin
User-agent: Launchmetrics
User-agent: LexisNexis
User-agent: Liana
User-agent: magpie-crawler
User-agent: Make/production
User-agent: MarketResearch.com
User-agent: MarketWatch
User-agent: MarketWise
User-agent: Markit Digital
User-agent: Mediatoolkitbot
User-agent: moduleQ
User-agent: MONITIO
User-agent: MoodleBot
User-agent: Moody's
User-agent: Moreover
User-agent: MORNINGSTAR
User-agent: MuckRack
User-agent: Netvibes
User-agent: news-api.org
User-agent: Newslitbot
User-agent: NewsNow
User-agent: Northern Light
User-agent: Opinion-tracker
User-agent: Orbis
User-agent: Opoint
User-agent: Paqlebot
User-agent: Press Monitor Europe
User-agent: PressEngineBot D
User-agent: PriberamBot
User-agent: QuoteMedia
User-agent: QWAM CONTENT INTELLIGENCE
User-agent: RankurBot
User-agent: RavenPack
User-agent: ReportLinker
User-agent: Research & Markets
User-agent: S&P Capital IQ
U
Machine-readable extras
llms.txt
absent
ai.txt
absent
Sitemap
absent
Feeds
https://www.courrierinternational.com/feed/al…
Schema types
WebSite, SearchAction, NewsMediaOrganization
Edge
Fastly
Final URL
https://www.courrierinternational.com/
HTML size
229 KB, 2,759 words of text

Compare it with a competitor head to head

Put courrierinternational.com next to another domain and see which one an answer engine can actually read. Both sides are measured the same way, so the difference is the finding.

Already measured: vs cgtrader.com · vs genetec.com · vs pulsz.com

Publish the score free

Sites that score well embed the badge. It links back to this live report, which re-measures on every scan.

AI access score badge
<a href="https://crawlcensus.com/site/courrierinternational.com"><img src="https://crawlcensus.com/badge/courrierinternational.com.svg" alt="AI access score for courrierinternational.com" width="150" height="20"></a>

Track changes on this domain free for one

Crawler policy is edited quietly. We re-scan monitored domains daily, keep the history, and email you the moment a crawler is blocked or unblocked, an llms.txt appears, or the score moves.

Free for one domain per address, and always free for domains you have verified you own. No newsletter, only change alerts.

Pro follows 25 domains, keeps 90 days of history, compares them against each other and exports the lot as CSV, for 19 dollars a month. That is the difference: measuring is free, watching a portfolio is not.