Cloudflare's 15 September AI Crawler Block: What UK Sites Do Now
From 15 September 2026, Cloudflare’s default settings block mixed-use AI crawlers on every page that carries advertising, for free-tier and new customers. That covers most small UK sites, including ones running display ads or affiliate links to fund content. Three crawler categories now exist: search, agent and training. Site owners must choose to block, allow or monetise each one before the deadline, or the new default chooses for them.
Key takeaways
- From 15 September 2026, Cloudflare defaults free-tier and new sites to blocking AI training and agent crawlers on any page that carries ads, leaving only search crawlers allowed.
- 36% of crawler activity now comes from mixed-use bots that blend search and training in one identity, such as Googlebot. That blend is exactly what forced the three-way split.
- Existing paying Cloudflare customers can override the new default in the dashboard. New customers, new sites and all free-tier domains get it automatically on 15 September.
- The alternative to blocking is Pay Per Use, which replaced 2025’s Pay Per Crawl: a minimum $0.01 lands only when an AI answer actually uses your content, not on every fetch.
- Block the wrong bots and you fund nothing but you also earn no citations from them. Allow the wrong bots and you may still earn nothing if your page is never the one cited.
- Our own crawler-access audits check this split before touching anything else, because a wrongly blocked bot can undo months of GEO work in one robots.txt line.
Why Cloudflare split crawlers into three categories
Cloudflare’s own data shows 36% of crawler activity now comes from mixed-use bots that will not declare a single purpose. Googlebot is the clearest example: one crawler identity feeds both Google Search indexing and Gemini’s training pipeline. Publishers who wanted search traffic had no way to keep that without also feeding training, short of blocking Googlebot outright and disappearing from Google entirely. Cloudflare’s fix separates every AI crawler into search, agent or training. A site can now allow the function it wants and block the one it does not, bot by bot rather than all or nothing.
What actually changes on 15 September
The new default applies to three groups: brand-new Cloudflare customers, new sites added by existing customers, and every domain still on the free tier. For all three, pages carrying ads switch to allowing search crawlers while blocking agent and training crawlers automatically. Existing paying customers keep whatever configuration they already have and can adjust it manually in the dashboard at any time. The practical effect: a small UK blog or affiliate site on Cloudflare’s free plan, which is most of them, gets a new crawler policy on 15 September. That happens whether anyone logs in to check it or not.
Block, allow, or monetise: the three options compared
| Option | What it means | Upside | Risk |
|---|---|---|---|
| Block training and agent bots | Only search crawlers reach the site | No free training data handed over; matches the new default | Zero chance of any future citation the model might have surfaced from this content |
| Allow everything | Search, agent and training bots all reach the site | Maximum chance of being sampled and cited across every engine | Content trains competing models for free unless Pay Per Use is enabled and paid |
| Enable Pay Per Use per bot | Search allowed, training/agent priced per successful retrieval | Minimum $0.01 lands only when your content is actually used in an answer | Coverage depends on which AI companies have signed on; not universal yet |
None of the three is correct for every site. A news publisher relying on ad revenue from human readers has a different calculation to a GEO client whose entire commercial goal is being cited by ChatGPT or Perplexity. The decision belongs on the same list as robots.txt review and WAF rules, not left to whatever the default becomes on 15 September.
robots.txt and WAF specifics to check before the deadline
Cloudflare’s dashboard change does not replace robots.txt. A site can be correctly configured in Cloudflare’s Bot Management panel and still carry an old Disallow: / line written years ago for an unrelated reason. That single line blocks every AI crawler regardless of the new categories. Check both layers: the Cloudflare AI Crawl Control settings per category, and the robots.txt file itself for blanket disallow rules that predate any GEO strategy. Free-tier customers can access these category toggles too, so cost is not the barrier; awareness is. WAF custom rules matter only if a site has manually blocked specific user agents in the past, which is worth an audit even on sites that never touched Cloudflare’s AI settings directly.
What we check in a crawler-access audit
We ran a crawler-access check on a UK ecommerce client last year and found GPTBot blocked by an old robots.txt rule. It had been aimed at an unrelated scraper years before anyone on the team had heard of GEO. Removing that single line was the fix, not a content rewrite. That same site later became the Garden Ornaments case study behind this site, which took organic traffic from 727 to 6,370 monthly visits in seven months using no net-new referring domains. None of that growth would have shown up in AI answers if the crawler blocking the citation from ever happening had stayed in place. A crawler-access audit is the cheapest fix in GEO, because it removes a block rather than adds a feature.
Matt’s Pick
For most small UK sites running ads and chasing AI visibility, leave search and agent crawlers allowed, since agent traffic carries real referral value. Treat training crawlers as the one genuine judgement call. Blocking training rarely costs a citation today, since citation-generating retrieval sits closer to the agent and search categories. Run the crawler-access check before 15 September rather than after, because the fix, when there is one, is usually a single line to remove.
Frequently asked questions
What happens on 15 September 2026?
Cloudflare defaults to blocking mixed-use AI crawlers on any page carrying ads, for free-tier and new customers.
Does this affect Googlebot?
Only the training/agent side. Cloudflare splits crawlers into search, agent and training, and Googlebot’s search function stays allowed by default.
Can I keep my current settings after 15 September?
Yes, if you are an existing paying Cloudflare customer. Free-tier and new sites get the new default automatically.
What is Pay Per Use, and is it different from Pay Per Crawl?
Yes. Pay Per Use replaced 2025’s Pay Per Crawl and pays a minimum $0.01 only when an AI answer actually uses your content, not on every fetch.
Should a small UK site just block every AI crawler?
No, not by default. Blocking training bots is usually safe, but blocking agent and search crawlers can remove the exact citations GEO work is trying to earn.
How do I check which bots are currently allowed on my site?
Open Cloudflare’s Bot Management or AI Crawl Control dashboard and review the search, agent and training categories per domain.
What does GeoLinks check in a crawler-access audit?
Which bots are blocked by robots.txt, WAF rules or Cloudflare defaults, and whether that setup helps or blocks the citations a site is paying to earn.
Related reading
- Cloudflare now pays publishers per AI citation covers the Pay Per Use pivot this policy builds on.
- Bots overtook humans: how to find out if AI can even see your brand explains why crawler access is the foundation everything else in GEO sits on.
- Why isn’t my brand in ChatGPT? 7 reasons lists crawler blocking as one of the most common, and most fixable, causes.
- llms.txt: the honest data on whether it works covers the other file AI crawlers are supposed to read, and how little it actually changes on its own.
Run a crawler-access audit before 15 September to see exactly which bots reach your site today.