SM/RSoftware Marketing Resource Subscribe

Blocking AI training no longer costs you Googlebot, and it still won't keep you out of AI answers

Cloudflare's new Disallow AI Training setting separates model training from search crawling, and quietly changes what the old Block option does to your search visibility.

Sienna McphersonSienna McphersonContributing writer
Sep 16, 2026 · 4 min read
X in f
A laptop open on a white desk running a code editor, with a second display behind it
Crawler rules live in robots.txt and a dashboard most marketers cannot open. Photo: Unsplash

Cloudflare has pulled AI training apart from search crawling, so a site can now refuse to feed Google's, Apple's and Microsoft's models while staying in their search indexes. The control is called Disallow AI Training, it took effect on September 15, and it quietly rewrites what the older blocking options do for anyone who set them months ago.

The knot it cuts is that the largest crawlers do two jobs on one visit — Cloudflare calls them mixed-use — so refusing the training half used to cost you the search half. Googlebot, Applebot and Bingbot were therefore carved out of the network's training blocks altogether.

The new setting, and the one that changed underneath it

Disallow AI Training writes a Disallow directive into your robots.txt and leans on the operator to respect it. Cloudflare has invented a label for the operators it believes will — Accountable — and Apple, Google and Microsoft all carry it, so their crawlers keep arriving for search. Training-only bots from Anthropic, Meta, OpenAI and Amazon are stopped under the same setting, at no cost in search, since all four crawl separately for indexing. Accountable status requires four things, offered or dated:

  • a way to refuse training via robots.txt or a comparable standard
  • an AI summaries opt-out, arranged direct today and routed through Cloudflare next year
  • URL-level reporting on what was handed to training, and on how it performed in search
  • a guarantee that refusing training carries no ranking penalty

The second change is retroactive. Block now means block. It used to spare mixed-use crawlers precisely to protect discoverability; it no longer does, so a site with that option selected is out of Google, Bing and Apple search. Sites already blocking training are being moved onto the new setting automatically, and two older options, Block AI Bots and Managed Robots.txt, are being retired.

It does not decide whether you appear in AI answers

This is where a software marketer can lose a quarter. At Google, the Cloudflare switch resolves to a robots.txt rule against Google-Extended, which governs model training and nothing else. Whether your pages surface in AI Overviews or AI Mode is decided on a different toggle inside Search Console, and Search Engine Journal notes that Google's documentation says neither lever touches the other. Flip the new control to keep your content out of AI, and your content stays in AI.

Bing is a wider gap. Microsoft is still building robots.txt support for a no-training preference and is aiming at early 2027, so the new setting conveys nothing there in the meantime. The one working lever is the NOARCHIVE meta tag, which per Bing's own documentation also stops Copilot linking to the page. Any software company chasing Copilot citations should leave it alone.

"Mixed-use crawlers were the hard part of the training question. AI Summaries are next."

That is Cloudflare's own account of what it has not solved yet; the summaries control it is building is a dial rather than a switch, aimed at early next year.

Publishers and software companies are not doing the same arithmetic

Cloudflare's presets for new domains are stricter for sites that sell advertising, on the logic that an impression only pays when a person loads the page. That reasoning does not port to a B2B software site, where the page is not inventory. A visit is worth a trial, a demo request or a docs read that ends in an API key — none of them priced per pageview.

Its own figures point the same way: fewer than one site in a hundred on the network turns search crawlers away, while 17% have something switched on against training. And it puts AI-search referrals at three to five times the conversion rate of conventional ones, against summary readers being over 40% likelier to stop there. Fewer arrivals, better ones — a trade that reads differently when you sell seats instead of impressions.

17%Cloudflare sites blocking AI training
<1%blocking search crawlers
3-5xAI-search referral conversion

There is a sharper version of this for developer tools. Training is the mechanism by which a model learns your CLI flags, your SDK method names and the exact wording of your error messages. Refusing it on your documentation is a decision not to be the answer when someone asks an assistant how to do the thing your product does, while a competitor who left training on becomes the default. Docs and marketing pages deserve different answers, and these controls apply per domain, which makes a docs subdomain the unit of decision.

What a software marketer should do this week

The awkward part is ownership. At most software companies nobody in marketing holds a Cloudflare login. Bot settings sit with infrastructure or security, who are being migrated onto new defaults on the strength of a decision once made about scrapers, not demand generation. The distribution surface is being reconfigured in a panel the people responsible for distribution cannot see.

  • Ask whoever owns Cloudflare what the Training, Search and Agent controls read today, and what they were migrated from
  • Confirm nobody has Block selected, which now costs you Googlebot
  • Check the Search Console AI setting separately — that governs AI Overviews and AI Mode, not the Cloudflare switch
  • Decide docs and marketing site independently, since the controls are per domain
  • If Copilot citations matter, do not reach for NOARCHIVE
What to do

Treat this as a distribution setting rather than a security one, and get marketing into the room before migrated defaults harden by neglect. The argument worth having is not whether models may train on your content — it is whether your documentation should be the exception.

AI crawlersCloudflarerobots.txtAI search
Share: X · LinkedIn · Facebook
Read next

One briefing, every Tuesday.

The week in software marketing: the news that matters, one unsponsored review, and the numbers behind both.

Free · Unsubscribe anytime