CrawlPact

Insight · Web infrastructure

Cloudflare's Disallow AI Training Setting: Stay Discoverable in Search Without Allowing Training

Cloudflare's Disallow AI Training keeps accountable mixed-use crawlers such as Googlebot for search while refusing training. How it differs from Block.

By CrawlPactPublished
Diagram: one mixed-use crawler does search and training. Disallow AI Training lets search continue and refuses training; Block stops both, so search access goes too. Disallow preserves accountable search access; Block can stop both.

On 15 September 2026 Cloudflare added a Disallow AI Training setting to its AI bot controls and changed what Block means. The new setting lets a site keep being crawled for search by operators Cloudflare calls “Accountable” — Apple, Google and Microsoft — while refusing AI training. Block, meanwhile, now applies to mixed-use crawlers too, so choosing it stops Applebot, Bingbot and Googlebot entirely, search included.

Key takeaway. If your goal is “stay in search, stay out of training”, the Cloudflare setting to evaluate is Disallow AI Training, not Block. Block is now the stronger, broader control — and for mixed-use crawlers it removes search access as well.

This is an analysis of Cloudflare’s change. For the vendor-neutral decision and the robots.txt groups that implement it anywhere, see blocking AI training while staying visible in AI search and how to block only AI training crawlers.

Why mixed-use crawlers forced a trade-off

Cloudflare classifies automated traffic by behaviour, and three behaviours are available as controls: Search (“crawling to build a search index”), Training (“crawling to train or fine-tune a model”) and Agent (user-directed agents acting for a person). A mixed-use crawler is “a single crawler doing both Search and Training”.

That is where a simple block fails. If the same crawler identity feeds a search index and a training pipeline, an edge block or a User-agent block cannot say “yes to search, no to training” — refuse one and you refuse the other. Cloudflare puts numbers on the tension from its own network: fewer than 1% of Cloudflare sites block Search bots, while 17% enable some mechanism to block training.

Operators that run separate crawlers do not create the problem. Cloudflare names Amazon, Anthropic, Meta and OpenAI as operators that “separate their Search and Training crawlers, so Cloudflare can block the Training crawler without affecting search.”

The four Training settings

Cloudflare applies Training, Search and Agent controls at the domain level. As Cloudflare describes them on 15 September 2026, the Training options now work like this:

Objective Setting What Cloudflare does Search impact Nature
Allow search and training Allow Allows crawlers, unless another setting or a WAF rule blocks them None Access
Keep search, refuse training Disallow AI Training Publishes a no-training preference in robots.txt; keeps Accountable mixed-use crawlers allowed for search; blocks every other training crawler Designed to preserve search for Accountable mixed-use crawlers Preference plus enforcement
Refuse crawlers on ad-funded pages Block on pages with ads Blocks crawlers, including mixed-use crawlers, only on pages detected to be serving an ad Mixed-use crawlers lose access to those pages, search included Enforcement
Refuse training crawlers completely Block Blocks all training crawlers, including mixed-use crawlers Applebot, Bingbot and Googlebot are stopped — search included Enforcement

Two details in Cloudflare’s post are easy to miss. Disallow AI Training exists only for Training, not for Search or Agent. And there is no “Disallow AI Training on pages with ads”, because the preference is published in robots.txt and the set of ad-serving pages “is too large and changes too frequently to enumerate” there.

Disallow AI Training is not just a robots.txt line

The name suggests a polite request, but the setting does two things at once:

  • For Accountable mixed-use crawlers it is a preference: Cloudflare’s Bot Preference Sync publishes the no-training rule — for example a Disallow for Google-Extended or Applebot-Extended — and the crawler stays allowed for search on the understanding that its operator honours the opt-out.
  • For every other training crawler it is enforcement: Cloudflare says those crawlers are blocked, “including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI”.

So “policy expression versus technical enforcement” still matters, but it splits by operator. The training opt-out for Googlebot’s content rests on Google honouring Google-Extended; the opt-out for a training-only crawler is a block at Cloudflare’s edge.

What “Accountable” means — and does not

Accountable is Cloudflare’s own designation. To qualify, an operator must meet or commit to:

  1. a way to opt out of AI training through robots.txt or a similar standard;
  2. a way to opt out of AI summaries, set with the operator now and through Cloudflare “next year”;
  3. URL-level visibility into which pages were made available for training, plus search metrics;
  4. assurance that opting out of training “will not affect traditional search results”.

Cloudflare says Apple, Google and Microsoft meet these through a mix of current capabilities and time-bound commitments. That is a vendor designation, not a certification, and it does not mean every crawler on the internet honours a no-training preference.

The per-operator detail matters, because the three do not work the same way today:

Operator Training opt-out today (per Cloudflare, 15 September 2026) What it means for Disallow AI Training
Google Disallow for Google-Extended in robots.txt Conveyed through robots.txt; Google states Google-Extended does not affect Search inclusion
Apple Disallow for Applebot-Extended in robots.txt Conveyed through robots.txt; Apple has stated disallowing training does not affect search ranking
Microsoft The NOARCHIVE meta tag; a robots.txt no-training mechanism is targeted for early 2027 Not yet conveyed to Bing through robots.txt — Cloudflare says so explicitly

Cloudflare’s own words on Bing: “Until that support launches, selecting Disallow AI Training will not automatically convey a no-training preference to Bing through robots.txt.” A site that wants a Bing training opt-out today needs Bing’s own mechanism as well.

What changed for existing sites

For most sites, nothing has to be done. Cloudflare’s post says current settings carry over, and domains that had previously chosen Block or Block on pages with ads for Training were migrated to Disallow AI Training — preserving their practical effect, since the old Block did not apply to mixed-use crawlers. Cloudflare also announced that “Block AI Bots” is being deprecated in favour of the Search, Training and Agent controls, and Managed robots.txt in favour of Bot Preference Sync.

The one decision that now has to be explicit: “If you want mixed-use crawlers gone entirely, you now have to say so.” That is Block, with its search consequence.

A documentation caution: as of 2 October 2026, Cloudflare’s “Block AI bots” documentation page (last updated 1 July 2026) predates the Disallow AI Training setting and describes the earlier plan for mixed-purpose crawlers. The 15 September announcement is the most specific public description of current behaviour; check both before relying on either.

What to verify after choosing a setting

  1. Fetch your live robots.txt from the public internet. Bot Preference Sync prepends its managed groups to any existing file; confirm the result is the policy you intended, not just what the dashboard says.
  2. Confirm Block is not set for Training unless you mean it. If it is, expect Googlebot, Bingbot and Applebot to be stopped.
  3. Watch search health in Google Search Console and Bing Webmaster Tools after the change.
  4. Cover Bing separately if a training opt-out for Bing matters to you before Microsoft ships robots.txt support.
  5. Write the decision down: which setting expresses the preference, which one enforces it, and why.

Where CrawlPact fits

CrawlPact audits the public crawler-policy signals a website declares — including the combined robots.txt Cloudflare serves — against a source-backed crawler registry (see the methodology). An audit shows whether Google-Extended or GPTBot is disallowed while Googlebot stays allowed. It cannot see your Cloudflare dashboard, edge blocks or WAF rules, and it cannot confirm that any operator honours a preference. For how Cloudflare’s managed rules combine with your own file, see the Cloudflare platform guide.

The takeaway

Cloudflare’s change makes the policy question more precise. For mixed-use crawlers, the choice is no longer “allow everything” or “block everything”: Disallow AI Training keeps Accountable search crawlers in while refusing training, and blocks the training crawlers that are not Accountable. Block now means what it says — including for search. Choose the narrower control when that is your policy, verify the public robots.txt, and remember that Bing is not covered through robots.txt until early 2027.

Sources

Claims about vendors, protocols and standards trace to these sources, last verified . Statements about CrawlPact's own systems are CrawlPact's first-party observations.

Check what your website currently declares to crawlers

CrawlPact audits the public crawler-policy signals a site exposes — robots.txt groups, robots meta and X-Robots-Tag directives, and related declarations — against a source-backed crawler registry. It does not guarantee crawler compliance, search ranking or AI citation, and it does not measure actual crawler traffic.

Audit a domain

Spotted an error or an outdated source? Report it through CrawlPact's corrections process.

Analytics preferences

CrawlPact uses optional Google Analytics and Microsoft Clarity on public pages to learn which content is useful. Clarity records clicks and scrolling (session replay), with anything you type masked. Neither runs in the app or admin areas, and CrawlPact works the same whether you accept or decline. See the privacy policy.