CrawlState Site Review

7 September 2026 at 00:48 UTC

theguardian.com

A short, dated report on one thing: what your website tells the machines that read it, and what those machines actually do when they arrive. Everything below was either read out of our record or measured directly. Nothing is estimated.

Review another site

How well AI reads theguardian.comAI Readability Score

40/ 100Weak

3 of 3 crawler identities we sent were served your homepage. An assistant can reach your page and would struggle to say anything specific about your business from it.

Can AI reach you30/30

Does the page come back when an AI asks for it?

  • Search and citation10/10crawlstate-search asked for your homepage and was served it (HTTP 200).
  • Agent10/10crawlstate-agent asked for your homepage and was served it (HTTP 200).
  • Training10/10crawlstate-train asked for your homepage and was served it (HTTP 200).
Can AI understand you10/30

Is there anything on the page a machine can actually read?

  • Structured data0/12Your page carries no structured data. An assistant has to guess your details from the wording of the page.This is the single biggest thing on this list. A short block of hidden text stating your name, address, hours and what you sell, in the format every assistant already reads.
  • Page title5/5Your title reads “Latest news, sport and opinion from the Guardian”.
  • Description5/5Your page has a meta description, which is often what gets quoted back.
  • Headings0/4Your page has no main heading at all.One main heading per page. More than one and nothing is the subject; none and the machine has to guess.
  • Readable without JavaScript0/4Only about 48 characters of text arrive in the page itself; the rest is assembled by JavaScript afterwards.Some crawlers never run JavaScript. If your words only appear after the page has run scripts, those crawlers see an empty page.
Can AI recommend you0/25

Can it state your address, hours and phone number to someone?

  • Named as a business0/8Nothing on your page says what kind of business this is in a form a machine reads.Declaring the business type is what gets you into 'where should I eat' style answers rather than only into searches for your name.
  • Address0/6Your address is not in a machine-readable form.An assistant will not send someone to an address it had to guess from a picture of a map.
  • Phone number0/6Your phone number is not in a machine-readable form.A tel: link or a phone field in your structured data. This is the one that turns an answer into a booking.
  • Opening hours0/5Your opening hours are not stated in a form a machine can read.'Are they open now' is one of the most common things people ask an assistant. Without this it cannot answer for you.
Does 15 September change it0/15

Is the answer above about to be decided for you?

  • 15 September exposure0/15Your site is served through Cloudflare. Cloudflare has said that from 15 September 2026 it blocks AI training crawlers and AI agents by default on pages that carry advertising, unless the site owner opts out.This is a setting in the Cloudflare dashboard, not a change to your website. It takes minutes once you know where to look — but nobody is going to email you about it.

Every line above is something your site actually returned when we asked it, just now. How much each one is worth is CrawlState’s judgement, not a universal measurement — the evidence is printed next to the points so you can disagree with the weighting and still use the findings.

Tell me if this changes

Your site’s answer to AI can change without you touching anything — a host default, a plugin update, 15 September. We re-check theguardian.com once a night and email you only when something actually moves. Free, and you can stop it in one click.

One — What your site tells AI crawlers todayThe policy you publish

Every AI company that bothers to look reads one small file on your website before it reads anything else. It sits at theguardian.com/robots.txt, and most site owners have never opened it. This is what yours says.

Names AI crawlersNot recorded

The record keeps the strongest signal it found on a site, and this was not it.

Licence or priceYes

You publish terms, a price, or a payment response. You are asking to be paid.

Content-SignalNot recorded

The record keeps the strongest signal it found on a site, and this was not it.

Record dated6 September 2026

Your site answered us with HTTP 200.

Two — What actually happens when they knockThe part nobody checks

Publishing a policy and having it obeyed are two different things. We asked your site for its own homepage four times: once as an ordinary web browser, then once each as the three kinds of AI crawler that matter to a business like yours.

Your site publishes terms asking AI companies to pay for your content. All three we tested were served the page in full, and none of them paid anything.

A person’s browserServed the pageHTTP 200

The control. This is what a customer sees, and what the crawlers below are compared against.

Search and citationcrawlstate-searchServed the pageHTTP 200

Reads your page to answer someone's question, and names your business in the answer. crawlstate-search asked for your homepage and got it, in full, with nothing in the way.

Agentcrawlstate-agentServed the pageHTTP 200

Fetches your page in the moment, because a person asked an assistant about you. crawlstate-agent asked for your homepage and got it, in full, with nothing in the way.

Trainingcrawlstate-trainServed the pageHTTP 200

Collects your pages to train a model. Nothing comes back to you. crawlstate-train asked for your homepage and got it, in full, with nothing in the way.

Measured 7 September 2026 at 00:48 UTC, from our servers, in a single pass. One request is a snapshot: a crawler let through today can be refused tomorrow, and a site can treat a crawler differently on a page deeper in.

Three — What changes on 15 September 2026A decision that may get made for you

From 15 September 2026, Cloudflare blocks AI training crawlers and AI agents by default on pages that carry advertising, unless the site owner opts out. Cloudflare sits in front of a large share of the web, often without the owner ever choosing it — it comes bundled with the hosting. If it sits in front of yours, that change lands on you whether or not you have an opinion about it.

Served throughcloudflare

Taken from the server header your site returned to us just now.

AffectedYes

Your site is served through Cloudflare, so the September default applies to it. If you want AI agents to keep reaching your pages, that is now something you have to say out loud in the Cloudflare dashboard.

Four — How you compareAgainst 347,935 other websites

We swept 449,070 domains on 6 September 2026. 347,935 of them answered. This is what they declared, and it is not a survey or a sample — it is the whole sweep.

Name an AI crawler and refuse it58,098 · 16.7%
Publish a licence or a price — you1,848 · 0.53%
Declare nothing at all286,854 · 82.4%

You are one of 1,848 sites on the entire web that ask to be paid — half of one per cent. Whatever else is true, you are not the crowd.

Five — What we recommendCharge — carefully

You already publish terms, which puts you in the smallest group on the web. The question now is whether anyone is honouring them.

  • Check the licence page actually loads. Terms that return a 404 read to a crawler as no terms at all.
  • Terms are a stated price, not a lock. 3 of the three crawlers we tested took the page without paying, and nothing on your side stopped them.
  • Keep the crawlers that cite you allowed. An answer that names your business and links to you is free distribution, and blocking it costs you customers.

Six — If you want this fixedWe can do it for you

Everything above is yours to act on, free, and your web person can do all of it. If you would rather it were simply done, this is what we charge.

Basic

$99 one payment

We set your AI policy properly and fix the September exposure. You get this report again afterwards, showing the difference.

Get this done — $99

Everything else

Let’s talk

Structured data written for your business, hours and prices and location made machine-readable, and the whole score above taken as far as it goes. Priced once we have seen the site, because no two are the same amount of work.

Start a conversation

Watch it

$49 / month

We re-run this review every night and email you the day something changes — including anything your host changes without telling you.

Watch theguardian.com

No subscription on the first two, no account needed, and nothing about this report changes if you never buy anything. The longer version.

Appendix — for whoever built the siteThe same findings, technically

Forward this section to your web person or agency. It is the raw form of everything above: what was requested, what came back, and which markup was present. Measured against https://theguardian.com/.

passSearch and citationuser-agent: crawlstate-search/1.0 — HTTP 200
passAgentuser-agent: crawlstate-agent/1.0 — HTTP 200
passTraininguser-agent: crawlstate-train/1.0 — HTTP 200
failStructured dataNo application/ld+json block found.
passPage title<title> length 48
passDescriptionmeta[name="description"] present
failHeadings<h1> count: 0
failReadable without JavaScriptText outside script/style/template/noscript: 48 chars
failNamed as a businessNo LocalBusiness-family @type found
failAddressNo streetAddress or addressLocality
failPhone numberNo telephone property or tel: link
failOpening hoursNo openingHours property
fail15 September exposureserver: cloudflare

Requests were sent from Cloudflare’s network with a plain GET, redirects followed, an eight second ceiling and no retries. Markup checks read the HTML as delivered, before any JavaScript runs — which is what several AI crawlers see.

How to read this. Section one comes from the CrawlState record, a dated sweep of 449,070 domains on 6 September 2026. Section two was measured live when this page was built, by asking your homepage for itself under four identities and recording what came back. We do not read your server logs, and a private arrangement made through a CDN is invisible to us, because from outside it is invisible to everyone.

Questions are answered at crawlstate.com/faq, and the method at crawlstate.com/data. Reply to the email this arrived in and a person will answer.