The data

The record behind
the research.

This is the page for anyone who wants to check our working. CrawlState is built on one dataset: what we find when we sweep the web asking every domain what it tells AI crawlers, and then what it actually does when one arrives. Everything else on this site comes out of it. Every figure below is a count from a dated sweep, not an estimate, not a model, and not bought in from anybody.

The record · swept from 6 September 2026

449,070 domains read.
347,935 answered.

347,935
answered the sweep
58,098
mention AI crawlers in robots.txt — 16.7%
1,848
publish a price or return HTTP 402 — 0.53%
101,135
did not answer at all
Declare nothing about AI286,854 · 82.4%
Mention AI crawlers in robots.txt58,098 · 16.7%
Publish a price or licence1,848 · 0.53%
Content-Signal preference only1,135 · 0.3%

Percentages are of the 347,935 domains that answered. The other 101,135 timed out, refused the connection, or never responded, and they are recorded as unreachable rather than folded into any of the categories above. Separately, 215,555 of the swept domains sit behind Cloudflare, which is the group the 15 September default change applies to.

Named publishers found in the sweep
SignalDomains
Publishes a licencemedium.com · theguardian.com · bild.de · seattletimes.com · drugs.com · bustle.com
Returns HTTP 402telegraph.co.uk · dokuwiki.org

Method

Two requests
per domain.

GET the root. GET robots.txt. An honest user agent with a contact email, and we follow redirects when a server sends them. We check the response headers for pricing or licence signals and record what robots.txt says. That is the complete method — there is nothing else behind it.

We do not resolve DNS, execute JavaScript, or parse meta tags. We have no access to anybody’s server logs. We read robots.txt to record what a site declares about crawler access, and we fetch the root page regardless of what it says, which is the only way to report on what a domain publishes. We never spoof another crawler’s identity.

The sweep records declarations. The second measurement, the one behind every site check, is a separate thing: we request the homepage three times, once under each of three named AI crawler identities, and write down the status code that comes back. That is how we can say that a domain declaring a block served all three of them a normal page, or that a domain declaring nothing refused all three. We report both readings side by side and never reconcile them into a single verdict, because the disagreement is the finding.

What one domain costs us
GET  https://example.com/            → status, headers
GET  https://example.com/robots.txt  → directives

recorded:
  crawler-price / crawler-charged headers
  Link: rel=license
  toll-vendor fingerprints
  robots.txt License:
  robots.txt Content-Signal:
  User-agent lines naming AI crawlers
CADENCE

One sweep done, a larger one running

The first sweep ran on 6 September 2026. A larger sweep is running now. The intent is daily, though the schedule is not established yet and we would rather say that than publish one we cannot keep.

BLIND SPOT

What we cannot see, we call undeclared

If a domain’s policy is only visible to cryptographically verified crawlers — Cloudflare pricing works this way — it is invisible to us and the domain is recorded as undeclared. Same for anything living in a terms-of-service page.

NO BACKFILL

Nothing is invented and nothing is adjusted

Every figure comes from a sweep. If a domain did not answer, it is not in the record for that date, and we do not carry a previous answer forward to fill the gap.

TWO ANSWERS

Declared and actual are separate fields

The sweep figures on this page count what domains declare. What a domain actually serves an AI crawler is measured per site, on request, and recorded separately — because on a great many domains the two disagree. We never let one stand in for the other.

15 SEPTEMBER 2026

Cloudflare changes its defaults.

From 15 September 2026, Cloudflare blocks AI training and agent crawlers by default on ad-supported pages unless the owner opts out. 215,555 of the domains in this sweep sit behind Cloudflare, so a large number of sites will change their effective answer on a single day without publishing anything themselves, and without their owners being told. We will re-run sweeps after that date to record how much of the record moves.

Am I affected?

What you can rely on

Every number on this site is a count from a sweep we ran, on a date we can name. We do not model, extrapolate, or buy figures in. Where we cannot see something we say so, and “undeclared” means “not visible in the signals we scan” rather than “has no policy”. That distinction matters enough to spell out every time it comes up.