On 15 September, Cloudflare blocks AI training crawlers by default on ad-supported pages.We are recording the before →

Crawl permissions API

Know who lets AI in, before you crawl.

CrawlState records which websites allow, block, or charge AI crawlers, and updates it every day. Ask about one domain or half a million in a single call, and stop paying for requests that were always going to be refused.

domains tracked
0Kdomains tracked
refresh cycle
24hrefresh cycle
archive starts
Sep 6archive starts
POST /v1/lookup_5 of 5
  • theguardian.com

    licence declared

    charge
  • medium.com

    licence declared

    charge
  • bild.de

    licence declared

    charge
  • reddit.com

    no AI directives

    unknown
  • shopify.com

    no AI directives

    unknown

real results, top-1000 sweep, 2026-09-06

theguardian.com/medium.com/bild.de/nytimes.com/reddit.com/stackoverflow.com/shopify.com/wikipedia.org/substack.com/arstechnica.com/telegraph.co.uk/npr.org/theguardian.com/medium.com/bild.de/nytimes.com/reddit.com/stackoverflow.com/shopify.com/wikipedia.org/substack.com/arstechnica.com/telegraph.co.uk/npr.org/

What we record

Four states, checked every day.

Read from robots.txt directives, RSL licence declarations, content signals and live HTTP responses. Every state carries a word and a glyph, never colour alone.

allow

Allow

No directive stops you. The request will be answered, and it is worth spending on.

block

Block

The site names your crawler and refuses it. Every request is money burned and one step nearer a ban.

charge

Charge

Access is for sale. The site publishes machine-readable licence terms, or answers with HTTP 402 and a price.

unknown

Unknown

No directive either way. We say so rather than guess, because a guess inside a crawl budget is worse than a gap.

The collector

One sweep of the web, every night.

Every domain gets an honest request with a named user agent, its robots.txt read, its licence declarations parsed, its response recorded. No spoofed crawler identities, no ignoring the rules we are in the business of reporting.

Then the whole thing is written to cold storage and never touched again, so today stays queryable years from now.

The archive

A permission is not a fact. It is a date.

Sites change their minds. A crawl-rights answer from last month is not stale data, it is wrong data, and it will send your crawler at doors that closed weeks ago.

We keep every daily snapshot, so you can ask what changed rather than what is, and reconcile a crawl that ran in March against the rules that applied in March.

GET /v1/changes?since=2026-09-05
{
  "since":   "2026-09-05",
  "changed": 1284,
  "results": [
    {
      "domain": "example.com",
      "from":   "unknown",
      "to":     "block",
      "seen":   "2026-09-06T02:14:00Z"
    }
  ]
}

Pricing

Priced by lookups, not seats.

Free

Kick the tyres.

$0forever

  • 1,000 lookups a month
  • Single-domain lookups
  • Current state
Start free

Indie

One person, one project.

$19per month

  • 25,000 lookups a month
  • Bulk endpoint, 1,000 a call
  • Daily change feed
  • Email support
Start building

Operator

Common

A crawler in production.

$199per month

  • 500,000 lookups a month
  • Bulk endpoint, 50,000 a call
  • Daily change feed
  • 90 days of history
Get an API key

Archive

The whole record.

Talk to us

  • Unlimited lookups
  • The full archive, back to launch
  • Daily snapshot exports
  • Trend and aggregate queries
Get in touch