Cloudflare's 15 September AI Crawler Block: What UK Sites Do Now

By , Co-founder, GeoLinks · · 7 min read
A UK website owner reviewing Cloudflare bot management settings on a laptop screen, home office, morning light
A UK website owner reviewing Cloudflare bot management settings on a laptop screen, home office, morning light

From 15 September 2026, Cloudflare’s default settings block mixed-use AI crawlers on every page that carries advertising, for free-tier and new customers. That covers most small UK sites, including ones running display ads or affiliate links to fund content. Three crawler categories now exist: search, agent and training. Site owners must choose to block, allow or monetise each one before the deadline, or the new default chooses for them.

Key takeaways

  • From 15 September 2026, Cloudflare defaults free-tier and new sites to blocking AI training and agent crawlers on any page that carries ads, leaving only search crawlers allowed.
  • 36% of crawler activity now comes from mixed-use bots that blend search and training in one identity, such as Googlebot. That blend is exactly what forced the three-way split.
  • Existing paying Cloudflare customers can override the new default in the dashboard. New customers, new sites and all free-tier domains get it automatically on 15 September.
  • The alternative to blocking is Pay Per Use, which replaced 2025’s Pay Per Crawl: a minimum $0.01 lands only when an AI answer actually uses your content, not on every fetch.
  • Block the wrong bots and you fund nothing but you also earn no citations from them. Allow the wrong bots and you may still earn nothing if your page is never the one cited.
  • Our own crawler-access audits check this split before touching anything else, because a wrongly blocked bot can undo months of GEO work in one robots.txt line.
Close-up of a UK marketer's hands adjusting AI bot category toggles in a Cloudflare dashboard, desk lamp lighting
Search, agent and training are now three separate switches, not one.

Why Cloudflare split crawlers into three categories

Cloudflare’s own data shows 36% of crawler activity now comes from mixed-use bots that will not declare a single purpose. Googlebot is the clearest example: one crawler identity feeds both Google Search indexing and Gemini’s training pipeline. Publishers who wanted search traffic had no way to keep that without also feeding training, short of blocking Googlebot outright and disappearing from Google entirely. Cloudflare’s fix separates every AI crawler into search, agent or training. A site can now allow the function it wants and block the one it does not, bot by bot rather than all or nothing.

What actually changes on 15 September

The new default applies to three groups: brand-new Cloudflare customers, new sites added by existing customers, and every domain still on the free tier. For all three, pages carrying ads switch to allowing search crawlers while blocking agent and training crawlers automatically. Existing paying customers keep whatever configuration they already have and can adjust it manually in the dashboard at any time. The practical effect: a small UK blog or affiliate site on Cloudflare’s free plan, which is most of them, gets a new crawler policy on 15 September. That happens whether anyone logs in to check it or not.

Two UK content strategists reviewing a printed list of AI bot names and a laptop showing crawl logs, bright afternoon office
Reviewing which bots actually reach the site is the first step before the deadline, not the last.

Block, allow, or monetise: the three options compared

OptionWhat it meansUpsideRisk
Block training and agent botsOnly search crawlers reach the siteNo free training data handed over; matches the new defaultZero chance of any future citation the model might have surfaced from this content
Allow everythingSearch, agent and training bots all reach the siteMaximum chance of being sampled and cited across every engineContent trains competing models for free unless Pay Per Use is enabled and paid
Enable Pay Per Use per botSearch allowed, training/agent priced per successful retrievalMinimum $0.01 lands only when your content is actually used in an answerCoverage depends on which AI companies have signed on; not universal yet

None of the three is correct for every site. A news publisher relying on ad revenue from human readers has a different calculation to a GEO client whose entire commercial goal is being cited by ChatGPT or Perplexity. The decision belongs on the same list as robots.txt review and WAF rules, not left to whatever the default becomes on 15 September.

robots.txt and WAF specifics to check before the deadline

Cloudflare’s dashboard change does not replace robots.txt. A site can be correctly configured in Cloudflare’s Bot Management panel and still carry an old Disallow: / line written years ago for an unrelated reason. That single line blocks every AI crawler regardless of the new categories. Check both layers: the Cloudflare AI Crawl Control settings per category, and the robots.txt file itself for blanket disallow rules that predate any GEO strategy. Free-tier customers can access these category toggles too, so cost is not the barrier; awareness is. WAF custom rules matter only if a site has manually blocked specific user agents in the past, which is worth an audit even on sites that never touched Cloudflare’s AI settings directly.

A developer's screen showing a robots.txt file being edited next to a Cloudflare WAF rules panel, evening desk setup
An old robots.txt line can silently override a correct Cloudflare setting.

What we check in a crawler-access audit

We ran a crawler-access check on a UK ecommerce client last year and found GPTBot blocked by an old robots.txt rule. It had been aimed at an unrelated scraper years before anyone on the team had heard of GEO. Removing that single line was the fix, not a content rewrite. That same site later became the Garden Ornaments case study behind this site, which took organic traffic from 727 to 6,370 monthly visits in seven months using no net-new referring domains. None of that growth would have shown up in AI answers if the crawler blocking the citation from ever happening had stayed in place. A crawler-access audit is the cheapest fix in GEO, because it removes a block rather than adds a feature.

Matt’s Pick

For most small UK sites running ads and chasing AI visibility, leave search and agent crawlers allowed, since agent traffic carries real referral value. Treat training crawlers as the one genuine judgement call. Blocking training rarely costs a citation today, since citation-generating retrieval sits closer to the agent and search categories. Run the crawler-access check before 15 September rather than after, because the fix, when there is one, is usually a single line to remove.

Frequently asked questions

What happens on 15 September 2026?

Cloudflare defaults to blocking mixed-use AI crawlers on any page carrying ads, for free-tier and new customers.

Does this affect Googlebot?

Only the training/agent side. Cloudflare splits crawlers into search, agent and training, and Googlebot’s search function stays allowed by default.

Can I keep my current settings after 15 September?

Yes, if you are an existing paying Cloudflare customer. Free-tier and new sites get the new default automatically.

What is Pay Per Use, and is it different from Pay Per Crawl?

Yes. Pay Per Use replaced 2025’s Pay Per Crawl and pays a minimum $0.01 only when an AI answer actually uses your content, not on every fetch.

Should a small UK site just block every AI crawler?

No, not by default. Blocking training bots is usually safe, but blocking agent and search crawlers can remove the exact citations GEO work is trying to earn.

How do I check which bots are currently allowed on my site?

Open Cloudflare’s Bot Management or AI Crawl Control dashboard and review the search, agent and training categories per domain.

Which bots are blocked by robots.txt, WAF rules or Cloudflare defaults, and whether that setup helps or blocks the citations a site is paying to earn.

Run a crawler-access audit before 15 September to see exactly which bots reach your site today.