CrawlState Site Review
7 September 2026 at 00:48 UTC
theguardian.com
A short, dated report on one thing: what your website tells the machines that read it, and what those machines actually do when they arrive. Everything below was either read out of our record or measured directly. Nothing is estimated.
How well AI reads theguardian.comAI Readability Score
3 of 3 crawler identities we sent were served your homepage. An assistant can reach your page and would struggle to say anything specific about your business from it.
Does the page come back when an AI asks for it?
- Search and citation10/10crawlstate-search asked for your homepage and was served it (HTTP 200).
- Agent10/10crawlstate-agent asked for your homepage and was served it (HTTP 200).
- Training10/10crawlstate-train asked for your homepage and was served it (HTTP 200).
Is there anything on the page a machine can actually read?
- Structured data0/12Your page carries no structured data. An assistant has to guess your details from the wording of the page.This is the single biggest thing on this list. A short block of hidden text stating your name, address, hours and what you sell, in the format every assistant already reads.
- Page title5/5Your title reads “Latest news, sport and opinion from the Guardian”.
- Description5/5Your page has a meta description, which is often what gets quoted back.
- Headings0/4Your page has no main heading at all.One main heading per page. More than one and nothing is the subject; none and the machine has to guess.
- Readable without JavaScript0/4Only about 48 characters of text arrive in the page itself; the rest is assembled by JavaScript afterwards.Some crawlers never run JavaScript. If your words only appear after the page has run scripts, those crawlers see an empty page.
Can it state your address, hours and phone number to someone?
- Named as a business0/8Nothing on your page says what kind of business this is in a form a machine reads.Declaring the business type is what gets you into 'where should I eat' style answers rather than only into searches for your name.
- Address0/6Your address is not in a machine-readable form.An assistant will not send someone to an address it had to guess from a picture of a map.
- Phone number0/6Your phone number is not in a machine-readable form.A tel: link or a phone field in your structured data. This is the one that turns an answer into a booking.
- Opening hours0/5Your opening hours are not stated in a form a machine can read.'Are they open now' is one of the most common things people ask an assistant. Without this it cannot answer for you.
Is the answer above about to be decided for you?
- 15 September exposure0/15Your site is served through Cloudflare. Cloudflare has said that from 15 September 2026 it blocks AI training crawlers and AI agents by default on pages that carry advertising, unless the site owner opts out.This is a setting in the Cloudflare dashboard, not a change to your website. It takes minutes once you know where to look — but nobody is going to email you about it.
Every line above is something your site actually returned when we asked it, just now. How much each one is worth is CrawlState’s judgement, not a universal measurement — the evidence is printed next to the points so you can disagree with the weighting and still use the findings.
One — What your site tells AI crawlers todayThe policy you publish
Every AI company that bothers to look reads one small file on your website before it reads anything else. It sits at theguardian.com/robots.txt, and most site owners have never opened it. This is what yours says.
The record keeps the strongest signal it found on a site, and this was not it.
You publish terms, a price, or a payment response. You are asking to be paid.
The record keeps the strongest signal it found on a site, and this was not it.
Your site answered us with HTTP 200.
Two — What actually happens when they knockThe part nobody checks
Publishing a policy and having it obeyed are two different things. We asked your site for its own homepage four times: once as an ordinary web browser, then once each as the three kinds of AI crawler that matter to a business like yours.
Your site publishes terms asking AI companies to pay for your content. All three we tested were served the page in full, and none of them paid anything.
The control. This is what a customer sees, and what the crawlers below are compared against.
Reads your page to answer someone's question, and names your business in the answer. crawlstate-search asked for your homepage and got it, in full, with nothing in the way.
Fetches your page in the moment, because a person asked an assistant about you. crawlstate-agent asked for your homepage and got it, in full, with nothing in the way.
Collects your pages to train a model. Nothing comes back to you. crawlstate-train asked for your homepage and got it, in full, with nothing in the way.
Measured 7 September 2026 at 00:48 UTC, from our servers, in a single pass. One request is a snapshot: a crawler let through today can be refused tomorrow, and a site can treat a crawler differently on a page deeper in.
Three — What changes on 15 September 2026A decision that may get made for you
From 15 September 2026, Cloudflare blocks AI training crawlers and AI agents by default on pages that carry advertising, unless the site owner opts out. Cloudflare sits in front of a large share of the web, often without the owner ever choosing it — it comes bundled with the hosting. If it sits in front of yours, that change lands on you whether or not you have an opinion about it.
Taken from the server header your site returned to us just now.
Your site is served through Cloudflare, so the September default applies to it. If you want AI agents to keep reaching your pages, that is now something you have to say out loud in the Cloudflare dashboard.
Four — How you compareAgainst 347,935 other websites
We swept 449,070 domains on 6 September 2026. 347,935 of them answered. This is what they declared, and it is not a survey or a sample — it is the whole sweep.
You are one of 1,848 sites on the entire web that ask to be paid — half of one per cent. Whatever else is true, you are not the crowd.
Five — What we recommendCharge — carefully
You already publish terms, which puts you in the smallest group on the web. The question now is whether anyone is honouring them.
- Check the licence page actually loads. Terms that return a 404 read to a crawler as no terms at all.
- Terms are a stated price, not a lock. 3 of the three crawlers we tested took the page without paying, and nothing on your side stopped them.
- Keep the crawlers that cite you allowed. An answer that names your business and links to you is free distribution, and blocking it costs you customers.
Six — If you want this fixedWe can do it for you
Everything above is yours to act on, free, and your web person can do all of it. If you would rather it were simply done, this is what we charge.
Basic
$99 one payment
We set your AI policy properly and fix the September exposure. You get this report again afterwards, showing the difference.
Get this done — $99Everything else
Let’s talk
Structured data written for your business, hours and prices and location made machine-readable, and the whole score above taken as far as it goes. Priced once we have seen the site, because no two are the same amount of work.
Start a conversationWatch it
$49 / month
We re-run this review every night and email you the day something changes — including anything your host changes without telling you.
Watch theguardian.comNo subscription on the first two, no account needed, and nothing about this report changes if you never buy anything. The longer version.
Appendix — for whoever built the siteThe same findings, technically
Forward this section to your web person or agency. It is the raw form of everything above: what was requested, what came back, and which markup was present. Measured against https://theguardian.com/.
Requests were sent from Cloudflare’s network with a plain GET, redirects followed, an eight second ceiling and no retries. Markup checks read the HTML as delivered, before any JavaScript runs — which is what several AI crawlers see.
How to read this. Section one comes from the CrawlState record, a dated sweep of 449,070 domains on 6 September 2026. Section two was measured live when this page was built, by asking your homepage for itself under four identities and recording what came back. We do not read your server logs, and a private arrangement made through a CDN is invisible to us, because from outside it is invisible to everyone.
Questions are answered at crawlstate.com/faq, and the method at crawlstate.com/data. Reply to the email this arrived in and a person will answer.