CrawlState Site Review
7 September 2026 at 00:48 UTC
nytimes.com
A short, dated report on one thing: what your website tells the machines that read it, and what those machines actually do when they arrive. Everything below was either read out of our record or measured directly. Nothing is estimated.
How well AI reads nytimes.comAI Readability Score
3 of 3 crawler identities we sent were served your homepage. An assistant can read your page, but it is missing details it needs to recommend you with any confidence.
Does the page come back when an AI asks for it?
- Search and citation10/10crawlstate-search asked for your homepage and was served it (HTTP 200).
- Agent10/10crawlstate-agent asked for your homepage and was served it (HTTP 200).
- Training10/10crawlstate-train asked for your homepage and was served it (HTTP 200).
Is there anything on the page a machine can actually read?
- Structured data12/12Your page carries structured data (WebSite, ImageObject, SearchAction, EntryPoint), which is the format assistants read first.
- Page title3/5Your title reads “The New York Times - Breaking News, US News, World News and Videos”, which is 66 characters and will be cut short.The title is the first thing read and often the only thing quoted. It should name the business and what it is, in under 65 characters.
- Description5/5Your page has a meta description, which is often what gets quoted back.
- Headings4/4Your page has exactly one main heading, which tells a reader what the page is about.
- Readable without JavaScript4/4There are about 5,794 characters of real text in the page as delivered.
Can it state your address, hours and phone number to someone?
- Named as a business0/8Nothing on your page says what kind of business this is in a form a machine reads.Declaring the business type is what gets you into 'where should I eat' style answers rather than only into searches for your name.
- Address0/6Your address is not in a machine-readable form.An assistant will not send someone to an address it had to guess from a picture of a map.
- Phone number0/6Your phone number is not in a machine-readable form.A tel: link or a phone field in your structured data. This is the one that turns an answer into a booking.
- Opening hours0/5Your opening hours are not stated in a form a machine can read.'Are they open now' is one of the most common things people ask an assistant. Without this it cannot answer for you.
Is the answer above about to be decided for you?
- 15 September exposure0/15Your site is served through Cloudflare. Cloudflare has said that from 15 September 2026 it blocks AI training crawlers and AI agents by default on pages that carry advertising, unless the site owner opts out.This is a setting in the Cloudflare dashboard, not a change to your website. It takes minutes once you know where to look — but nobody is going to email you about it.
Every line above is something your site actually returned when we asked it, just now. How much each one is worth is CrawlState’s judgement, not a universal measurement — the evidence is printed next to the points so you can disagree with the weighting and still use the findings.
One — What your site tells AI crawlers todayThe policy you publish
Every AI company that bothers to look reads one small file on your website before it reads anything else. It sits at nytimes.com/robots.txt, and most site owners have never opened it. This is what yours says.
Your robots.txt names 8 AI crawler(s), and tells them not to read the site.
The record keeps the strongest signal it found on a site, and this was not it.
The record keeps the strongest signal it found on a site, and this was not it.
Your site answered us with HTTP 200.
Two — What actually happens when they knockThe part nobody checks
Publishing a policy and having it obeyed are two different things. We asked your site for its own homepage four times: once as an ordinary web browser, then once each as the three kinds of AI crawler that matter to a business like yours.
Your robots.txt asks AI crawlers not to take your content. All three we tested were served the page normally.
The control. This is what a customer sees, and what the crawlers below are compared against.
Reads your page to answer someone's question, and names your business in the answer. crawlstate-search asked for your homepage and got it, in full, with nothing in the way.
Fetches your page in the moment, because a person asked an assistant about you. crawlstate-agent asked for your homepage and got it, in full, with nothing in the way.
Collects your pages to train a model. Nothing comes back to you. crawlstate-train asked for your homepage and got it, in full, with nothing in the way.
Measured 7 September 2026 at 00:48 UTC, from our servers, in a single pass. One request is a snapshot: a crawler let through today can be refused tomorrow, and a site can treat a crawler differently on a page deeper in.
Three — What changes on 15 September 2026A decision that may get made for you
From 15 September 2026, Cloudflare blocks AI training crawlers and AI agents by default on pages that carry advertising, unless the site owner opts out. Cloudflare sits in front of a large share of the web, often without the owner ever choosing it — it comes bundled with the hosting. If it sits in front of yours, that change lands on you whether or not you have an opinion about it.
Taken from the server header your site returned to us just now.
Your site is served through Cloudflare, so the September default applies to it. If you want AI agents to keep reaching your pages, that is now something you have to say out loud in the Cloudflare dashboard.
Four — How you compareAgainst 347,935 other websites
We swept 449,070 domains on 6 September 2026. 347,935 of them answered. This is what they declared, and it is not a survey or a sample — it is the whole sweep.
You are in the 16.7% that names an AI crawler and refuses it. Most sites never get that far.
Five — What we recommendBlock, and mean it
You have already decided you do not want AI crawlers taking your pages. That decision is written in a file they are choosing not to read.
- Move the refusal from robots.txt to your hosting. A rule at the edge that answers named AI crawlers with a 403 is obeyed whether or not the crawler wants to obey it. Your host or web person can add it in an afternoon.
- Do not block the crawlers that cite you. When an AI answer names your business and links to your site, that is distribution you did not pay for. Refusing it is the one change here that can lose you bookings.
- Charging is almost certainly not worth it at your size. Licence deals are priced off traffic, and a site with a few thousand visits a month earns pennies a year. Charging is a publisher’s move, not a hotel’s.
Six — If you want this fixedWe can do it for you
Everything above is yours to act on, free, and your web person can do all of it. If you would rather it were simply done, this is what we charge.
Basic
$99 one payment
We set your AI policy properly and fix the September exposure. You get this report again afterwards, showing the difference.
Get this done — $99Everything else
Let’s talk
Structured data written for your business, hours and prices and location made machine-readable, and the whole score above taken as far as it goes. Priced once we have seen the site, because no two are the same amount of work.
Start a conversationWatch it
$49 / month
We re-run this review every night and email you the day something changes — including anything your host changes without telling you.
Watch nytimes.comNo subscription on the first two, no account needed, and nothing about this report changes if you never buy anything. The longer version.
Appendix — for whoever built the siteThe same findings, technically
Forward this section to your web person or agency. It is the raw form of everything above: what was requested, what came back, and which markup was present. Measured against https://nytimes.com/.
Requests were sent from Cloudflare’s network with a plain GET, redirects followed, an eight second ceiling and no retries. Markup checks read the HTML as delivered, before any JavaScript runs — which is what several AI crawlers see.
How to read this. Section one comes from the CrawlState record, a dated sweep of 449,070 domains on 6 September 2026. Section two was measured live when this page was built, by asking your homepage for itself under four identities and recording what came back. We do not read your server logs, and a private arrangement made through a CDN is invisible to us, because from outside it is invisible to everyone.
Questions are answered at crawlstate.com/faq, and the method at crawlstate.com/data. Reply to the email this arrived in and a person will answer.