Five weeks ago I published a piece warning that 33 of 140 healthcare membership organisations might stop appearing in Google on 15 September. The argument was that Cloudflare was about to start treating crawlers that do more than one job under whichever rule is most restrictive — and because Googlebot crawls both for Search and for AI training, an old “Block AI bots” setting might quietly start meaning “don’t index us either”.

The date has now passed. So this is the follow-up, and it is not the one I expected to write.

Two things need correcting. One is about the world. The other is about my own homework, and I am going to deal with that one first, because it is the more embarrassing of the two.


The correction: the figure of 33 was wrong

The number I published was wrong. Not rounded, not approximate — wrong, and produced by a mistake I should have caught before publishing.

Here is exactly what happened.

I ran the check twice. The first run, on 31 July, covered 66 organisations, of which 63 produced usable results. The second run, on 10 August, extended it to 145 organisations, of which 138 were usable. The article was written after the second run and quoted “140 organisations”, which was a loose description of the 145 scanned and 138 usable.

But the figure of 33 did not come from that second run at all. It came from the first one. In the 63-site July batch, 33 organisations were blocking training crawlers at the server. That is all that number ever meant: blocks training crawlers. It said nothing about Cloudflare, and nothing about whether Google could still reach the site.

By the time I wrote the article, I had defined the at-risk group much more tightly — behind Cloudflare, currently refusing training crawlers, and currently serving Googlebot normally. I then reported the old number against the new definition, from the wrong sample, and gave it a headline.

Two different runs, two different definitions, one number carried across between them. It is a boring error and it is entirely mine.

What the figure should have been

Going back to the 10 August data — 138 usable organisations — and applying the definition the article actually states:

Organisations scanned145
Usable after discarding failed and contaminated scans138
Behind Cloudflare50
Behind Cloudflare and blocking training crawlers32
…of those, still serving Googlebot normally31

So on the article’s own literal wording, the answer was 31, not 33. Close enough that nobody would have noticed — which is precisely why it is worth saying out loud. A number can be nearly right and still be reached by a method that does not work.

But 31 is not the number I would defend either, and here is why.

Of those 31, seventeen were not refusing training crawlers as a deliberate AI policy at all. They were refusing everything that was not a browser — search crawlers, assistant crawlers, training crawlers, the lot. That is a blanket bot rule or an over-enthusiastic security plugin, not somebody pressing “Block AI bots”. Cloudflare’s change acts on a training-specific signal, so those seventeen were never in scope for the thing I was warning about. They have a different problem, with a different fix.

Strip them out and you are left with 14 organisations of 138 — about one in ten — that genuinely matched the pattern I described: behind Cloudflare, blocking training crawlers specifically, with Google reaching them normally.

Fourteen. Not thirty-three. The article’s headline overstated the affected group by more than double.

For completeness, two neighbouring figures from the same data, since the Cloudflare-only framing may have been too narrow in the first place: 36 organisations were behind some edge vendor and blocking training crawlers, and 46 were blocking at least one AI assistant outright, today, with no deadline involved. That last one was correct in the original article and remains the more useful finding.


What actually happened on 15 September

Cloudflare shipped. But not the change I described.

On 15 September, Cloudflare published a post introducing a new setting called Disallow AI Training, alongside a designation it calls Accountable for crawler operators that meet a set of conditions — an opt-out mechanism for training, an opt-out for AI summaries, URL-level visibility, and an assurance that opting out of training will not affect search rankings.

Apple, Google and Microsoft all qualify. Which means Applebot, Bingbot and Googlebot are Accountable crawlers, and under Disallow AI Training they keep crawling your site for search while the no-training preference is published on your behalf.

The part that matters most for anyone who read my August article is the migration table. Cloudflare moved existing customers across automatically:

  • Domains with the legacy Block AI Bots setting turned on were migrated to Search: Allow, Training: Disallow AI Training.
  • Domains that had configured the granular controls with Training set to Block, or Block on pages with ads, were also migrated to Disallow AI Training.

Cloudflare’s own summary of what customers needed to do was: “Nothing, in almost every case. Your current settings carry over on their own.”

So the organisations I was worried about were not walked off a cliff. They were quietly moved onto a setting that does the thing they actually wanted in the first place — refuse training, stay in Google. The tradeoff I warned about was real, Cloudflare confirmed it was real, and then Cloudflare removed it.

The risk has not vanished entirely. It has moved. From 15 September, selecting Block for Training now does apply to mixed-use crawlers, so anyone who deliberately chooses Block will stop Applebot, Bingbot and Googlebot reaching their site — search included. That is now an active choice somebody has to make, rather than a consequence of an old toggle nobody remembered pressing. That is a much better place for it to sit.

It also makes the underlying decision easier to take on its merits, which is the conversation I would rather organisations were having anyway. I wrote about that separately in Should Your Association Block AI Crawlers? — and the short version is that training, retrieval and user-triggered access are three different things with three different sets of consequences, and it is entirely coherent to say no to the first while saying yes to the other two.


So what did I get wrong?

Three things, in descending order of how much they bother me.

The number. Covered above. A figure from one sample, reported against a definition from another. There is no interesting lesson in it — it is just a mistake, and the fix is to re-derive every published figure from the data it claims to come from before publishing, not after someone asks.

The prediction. I wrote that this was five weeks away and that organisations should get into the dashboard before the 15th. What I did not seriously consider was that Cloudflare would spend those five weeks negotiating with Apple, Google and Microsoft and ship a mitigation on the day. In hindsight the signal was there — the July announcement was framed as separating Search, Training and Agent precisely so people would not have to choose. A company that had just built three separate controls to avoid a blunt tradeoff was unlikely to then reintroduce the tradeoff and leave it there.

I read a vendor’s roadmap as a threat when it was closer to a work in progress. That is a specific kind of error and I would like to not repeat it.

The framing. “A quarter may lose Google” was the most alarming true-ish sentence available at the time. It was defensible on the definition I had written down. It was not defensible on the data. When the number and the headline disagree, the headline is the one that is wrong, and I should have gone back to the data rather than forward to the publish button.

To be clear about what I would still defend: the mechanism was real, the affected pattern existed, and fourteen organisations out of 138 were genuinely sitting on a setting that nobody in the building had chosen. That is worth knowing. It just is not a quarter of them, and nothing bad happened to them in the end.


What the fourteen had in common

I want to be careful here, because this is the section where I would most like to tell you a satisfying story and I cannot fully evidence one.

I have not re-run the scan. The toolkit needs to run from a server with ordinary outbound access, and the environment this follow-up was prepared in could not reach the sites, so a true like-for-like re-scan against the same 138 origins is still outstanding. I would rather publish that gap than fill it with an estimate. When I re-run it, I will update this piece with dated before-and-after figures.

What I can tell you is what the fourteen looked like in August, which is interesting in its own right.

All fourteen were WordPress sites. Every single one. Not because WordPress causes this — it does not — but because WordPress is what membership organisations run, and because the combination of a WordPress site, a security plugin, and a Cloudflare account opened by a third party is the standard architecture in this sector. Three systems, three sets of settings, and frequently three different people, none of whom is still in post.

None of the fourteen were blocking retrieval crawlers. That is what separated them from the seventeen I excluded. Their configuration was specific: training blocked, everything else allowed. Somebody, at some point, made a deliberate and sensible decision. It just was not written down anywhere, and it was not revisited when the ground moved.


The lesson, which was never really about Cloudflare

Here is the thing I would still say, and it survived the correction intact.

Nobody owns the settings.

The organisations in that group had not done anything wrong. Someone made a reasonable decision quickly, with the options available at the time, and then left. The Cloudflare account was opened by an agency in 2021. The notification about the change went to an email address belonging to a person who has not worked there for three years. Nothing looked broken, because nothing was broken — right up until the moment it might have been.

This time it ended well, and it ended well for reasons entirely outside anyone’s control. Cloudflare negotiated a better outcome and migrated everybody onto it automatically. That was luck, not readiness.

The next change will not necessarily come with a vendor who spends the summer fixing it on your behalf. And the reason this was frightening in August was never the specific setting. It was that a great many organisations could not answer a simple question: what do our settings currently say, and who decided that?

If you cannot answer that today, the answer is not to learn about Cloudflare. It is to find out who has the logins, write down what the current position is, and put a date in the calendar to look again. That is unglamorous, it is not a project, and it is worth more than any AI strategy you will be sold this year.

For what it is worth, Cloudflare’s own figures say less than 1% of sites on its network block search crawlers, while 17% use some mechanism to block training. Most people want to be found. They just want to be found on their own terms — and the only way to be sure you are is to know what you have currently agreed to.


If you would like me to check

The offer from the original article stands, and it is still free. Send me your domain and I will tell you what your site currently says to search crawlers, assistant crawlers and training crawlers, and whether any of it looks unintentional. It takes a few minutes and there is no pitch attached.

If you would like the fuller picture, our Website Check-Up is £400 and covers the technical condition of your site, accessibility and compliance, sustainability, usability, and what AI assistants currently say about your organisation. You will have the report within 48 hours of giving us access, usually within 24, and if you go on to work with us we deduct the £400 from that work.

If the AI question is the one that matters most — because accurate guidance is central to what you do — the AI Visibility Check is £450 and examines it properly on its own, or £750 for both together.

And if the real problem is that nobody has time to keep an eye on any of this, that is what the Website Care Plan is for. It starts at £400 per month, there is no minimum term, unused hours roll over for up to twelve months, and monthly AI visibility monitoring is included at no extra cost. Hosting, if you want it, is separate at £95–£240 +VAT per month.

Or simply book a call and we will talk it through.


A note on the research

The figures in this piece come from my own check of healthcare membership organisations carried out on 31 July and 10 August 2026. 145 organisations were scanned; 138 produced usable results after discarding scans that failed outright or showed signs of contamination from my own request rate. The original article described this as “140”, which was loose — the precise figures are above.

For each organisation I requested the homepage and the contact page as each of twenty documented crawlers, and compared every response against an ordinary browser request made moments earlier. Blocks that robots.txt explicitly permits are the ones counted here, because those are the blocks nobody chose.

Three limitations, all of which applied to the original article and still apply:

  • I cannot tell from outside what causes a block. Cloudflare, a WAF, a security plugin or a hand-written rule all look identical from the far end of an HTTP request. I can see the symptom, not the cause.
  • A single origin at a single moment. Cloudflare’s bot scoring is behavioural, so the same crawler can pass one day and be challenged the next.
  • Googlebot’s own agent is untestable from outside. Google’s user-triggered fetcher sends a standard Chrome user agent and is identifiable only by IP range. It is reported as unknown, never as passing.

The re-scan is outstanding. Every figure above describes August, not today. The like-for-like re-run against the same 138 origins has not been completed, and nothing in this piece should be read as a measurement of what changed after 15 September. When I have re-run it I will update this article with dated figures and say so at the top.

Sources

All URLs above were fetched and confirmed working on 16 September 2026.

Read related articles

icon a person with a speech bubble
Our thoughts

Your Members Finished the Course. Can They Prove It?

By Chris Plummer | | Education, Engagement
Read article
Icon opposite arrows moving through a magnifying glass
Our thoughts

We checked 140 health organisations. A quarter may lose Google on 15 September

By Chris Plummer | | AI, Development
Read article
icon a person with a tick mark speech bubble
Our thoughts

Who Actually Owns Your Website? The Handover Questions Nobody Asks

By Chris Plummer | | More Time To
Read article

Book a free no-obligation call

Are you ready to offload the hassle of managing your website and reclaim your time? Share your challenges, and let’s see how we can help.