Publishers using Cloudflare no longer have to choose between allowing Googlebot to crawl for Search and expressing an objection to Google using their content for AI training.
Cloudflare’s new Disallow AI Training setting is designed specifically to separate those purposes. The control publishes the appropriate no-training preferences through robots.txt while allowing qualifying mixed-use crawlers such as Googlebot and Applebot to continue accessing the site for traditional search.
The change, announced in Cloudflare’s September 15 technical update and detailed by Search Engine Journal, resolves an awkward problem created by crawlers that serve more than one purpose. But the solution is not symmetrical across Google, Apple and Microsoft yet.
Google and Apple already expose dedicated training opt-outs. Bing does not yet support the equivalent preference through robots.txt, leaving publishers with a temporary compromise: Microsoft’s current NOARCHIVE mechanism can opt content out of training, but it also prevents links to that content from appearing in Bing Chat and Copilot.
Cloudflare is separating “disallow training” from “block the crawler”
The most important technical distinction in the update is between a preference and a network block.
Cloudflare now classifies crawler behavior into three categories: Search, Training and Agent. A single crawler can perform more than one of those functions.
That creates a problem when one user agent is used for both Search and Training. Blocking the crawler at the network level removes both behaviors, even if the publisher only objects to training.
Disallow AI Training is Cloudflare’s answer to that problem.
Disallow AI Training keeps accountable mixed-use crawlers available for search
When the setting is selected, Cloudflare’s Bot Preference Sync publishes the relevant no-training directives in robots.txt.
Mixed-use crawlers from operators Cloudflare designates as “Accountable” can continue crawling for Search while honoring the publisher’s training preference.
Other training crawlers are blocked.
Cloudflare says this includes training-only crawlers operated by Amazon, Anthropic, Meta and OpenAI. Because those companies separate training crawlers from other relevant crawler functions, Cloudflare can block the training component without creating the same search-discoverability conflict.
Choosing Block now means exactly what it says
Publishers need to pay close attention to the neighboring Block option.
Cloudflare says Block and Block on pages with ads now apply to mixed-use crawlers as well. If a publisher chooses Block for the relevant category, Googlebot, Applebot and Bingbot can be stopped entirely.
That includes crawling performed for search.
The operational rule is therefore simple but consequential: use Disallow AI Training when the objective is to object to model training while preserving search access; use Block only when the objective is to stop the crawler itself under the selected control.
This changes the tradeoff Cloudflare warned about earlier this year
The September implementation is more nuanced than the policy Cloudflare described during the summer.
Cloudflare had previously warned that its stricter treatment of mixed-use crawlers could cause a site blocking Training to block Googlebot, Applebot and Bingbot because those crawlers were classified as performing both search and training functions.
The new Disallow AI Training option creates a separate path for cooperating operators.
Instead of enforcing the training objection by denying the mixed-use crawler access, Cloudflare can publish a machine-readable preference that the operator agrees to respect.
Cloudflare calls these operators “Accountable”
The new system depends on crawler operators cooperating with publisher preferences.
Cloudflare has therefore introduced an Accountable designation with four requirements that an operator must meet or commit to meeting within a specified timeframe.
Operators need a mechanism for publishers to opt out of AI training, a mechanism for controlling AI-summary use, URL-level transparency into which pages were made available for training together with search visibility metrics, and an assurance that opting out of training will not damage traditional search results.
Cloudflare says Apple, Google and Microsoft qualify under those criteria through a combination of current capabilities and time-bound commitments.
Google’s training opt-out uses Google-Extended
For Google, the mechanism is already familiar to technical SEOs.
Cloudflare publishes a Disallow rule for Google-Extended, Google’s control token for restricting specified uses of site content by Gemini models and related training or grounding contexts covered by that token.
Crucially, Google-Extended is not Googlebot.
Google has stated that restricting Google-Extended does not affect a site’s inclusion or ranking in traditional Google Search, and Cloudflare repeats that assurance in its announcement.
Googlebot can therefore keep crawling for traditional Search
This is the core SEO benefit of the new Cloudflare setting.
A publisher can continue allowing Googlebot to retrieve pages for Search while Bot Preference Sync expresses a separate objection through Google-Extended.
The site does not need to block Googlebot at the edge merely to communicate its AI-training preference.
For publishers whose organic search traffic remains economically important, that separation removes a potentially expensive technical conflict.
Google-Extended does not control AI Overviews or AI Mode
The distinction becomes more complicated when Google’s generative search features enter the discussion.
Cloudflare and Search Engine Journal both emphasize that the training opt-out should not be confused with controls over Google’s AI-generated search experiences.
Disallowing Google-Extended does not determine whether pages can appear in AI Overviews or AI Mode.
Those features belong to Google Search and are governed separately.
A publisher therefore cannot assume that “no AI training” means “do not use my content anywhere in AI-powered Search.”
Training and generative search visibility are separate policy decisions
This separation matters because publishers may want different outcomes.
A news organization could object to its journalism being used for model training while still wanting citations, links and visibility inside AI-powered search results.
Another publisher may want neither.
A third may allow training but restrict snippets or generative presentation for other commercial reasons.
The technical ecosystem is slowly moving toward controls that let publishers express those choices independently instead of forcing one all-or-nothing robots.txt decision.
Google is also promising more URL-level transparency
Cloudflare says Google is working on additional URL-level transparency tools related to Google-Extended and expects them to launch in the coming weeks.
The stated goal is to let site owners see which pages have been made available under the Google-Extended policy.
That is important because a robots.txt preference is more useful when publishers can verify its practical scope.
Today, site owners can publish a directive. Better URL-level reporting would help them inspect how that preference maps to individual content.
Apple already has Applebot-Extended
Apple’s architecture is similar in principle.
Publishers can use Applebot-Extended in robots.txt to communicate that content should not be used for training Apple’s generative models.
Cloudflare says Apple has stated that this training preference does not affect search ranking.
Apple also uses separate mechanisms for generative summaries: Cloudflare notes that the nosnippet directive can be used to express preferences regarding AI-generated answers to broad knowledge questions in Siri and Search.
Again, training control and generative presentation are not the same thing.
Apple’s URL-level inspection tooling is still in development
Cloudflare says Apple does not yet provide the URL-level inspection capability contemplated by the Accountable framework.
The company has nevertheless shared plans with Cloudflare for a solution targeted for next year.
This illustrates how Cloudflare is using the Accountable label.
It does not mean every operator has every transparency feature today. Some qualify because they have made concrete commitments with delivery timeframes.
Bing is the awkward exception
Microsoft’s Bingbot creates the most important limitation in the new system.
Bing does not currently support a robots.txt directive equivalent to Google-Extended or Applebot-Extended for AI-training opt-out.
Microsoft is building support for a site-level no-training preference through robots.txt, according to Cloudflare, with delivery targeted for early 2027.
Until that exists, selecting Disallow AI Training in Cloudflare does not automatically communicate a no-training instruction to Bing through robots.txt.
Microsoft currently relies on NOARCHIVE
Publishers that want to express a training opt-out to Microsoft today need to use the NOARCHIVE meta tag.
Microsoft has said that NOARCHIVE prevents the marked content from being used to train its generative AI models without affecting traditional search ranking.
That sounds similar to the Google and Apple position until the effect on conversational products is considered.
Bing documentation also says content marked NOARCHIVE will not be linked in Bing Chat and Copilot.
Bing therefore still creates a visibility tradeoff
This is the temporary compromise publishers cannot fully avoid through Cloudflare’s new robots.txt setting.
A site can tell Microsoft not to use content for AI training through NOARCHIVE while remaining eligible for conventional search ranking.
But the same choice removes links to that content from Bing Chat and Copilot.
For publishers pursuing AI referral traffic, that matters.
The cost of the training opt-out is not classic SEO visibility; it is conversational citation and link visibility inside Microsoft’s AI experiences.
That makes Bing strategically different from Google today
With Google, a publisher can disallow Google-Extended without automatically removing itself from traditional Search or using that token as a control for AI Overviews and AI Mode.
With Bing’s current NOARCHIVE mechanism, the training preference carries an additional consequence for Chat and Copilot links.
The two opt-outs therefore should not be treated as equivalent simply because both can reduce training use.
Publishers need to evaluate the downstream visibility effect platform by platform.
Microsoft plans to close the gap in early 2027
Cloudflare says Microsoft is developing support for a robots.txt no-training preference at the domain or site level.
The target is early 2027.
If implemented as described, that should give publishers a cleaner way to communicate training preferences without relying solely on NOARCHIVE.
Until then, Cloudflare customers need to understand that Disallow AI Training does not magically create a Bing directive that Microsoft does not yet support.
Cloudflare also points to Bing removal tools
For Cloudflare customers seeking stronger restrictions today, the company mentions Microsoft’s Block URLs and Content Removal tools alongside NOARCHIVE.
Those tools solve different problems and should not be treated as a universal replacement for a standardized training preference.
The broader point is that Microsoft’s current control surface remains more fragmented than Google’s or Apple’s for this specific use case.
That fragmentation is exactly what Cloudflare says it wants the industry to eliminate.
Bot Preference Sync is becoming the control layer
Cloudflare’s implementation depends on Bot Preference Sync, which connects a customer’s dashboard choices to the preferences published in robots.txt.
The system can prepend applicable crawler directives while preserving existing rules.
This reduces the maintenance burden for publishers that would otherwise need to monitor every crawler operator’s changing user agents and policy mechanisms manually.
Cloudflare’s role is not merely to write robots.txt. It also classifies crawler behavior and can enforce blocks at the network layer when appropriate.
Managed Robots.txt is being deprecated
As part of the transition, Cloudflare says its older Managed Robots.txt feature will be deprecated in favor of Bot Preference Sync.
The legacy Block AI Bots control is also giving way to the more granular Search, Training and Agent categories.
Existing customers will generally be migrated automatically.
Cloudflare says previous Training selections of Block or Block on pages with ads will move to Disallow AI Training under the new definitions, preserving the intended training restriction without unnecessarily blocking accountable mixed-use search crawlers.
Most existing customers do not need to change anything
Cloudflare says the migration is designed to carry current preferences forward in almost every case.
Sites that previously used the older Block AI Bots setting will be mapped into the new controls according to their existing configuration.
But technical SEO teams should still verify the result.
A crawler-control migration can affect organic discoverability if the wrong category ends up blocked, and the meaning of Block is now more consequential for mixed-use crawlers than it was previously.
New ad-supported domains get a more restrictive training preset
Cloudflare is also changing onboarding recommendations for new domains.
Sites that monetize pages with advertising are offered a preset that keeps Search allowed while setting Training to Disallow AI Training and restricting Agent behavior on pages where ads are detected.
Cloudflare’s reasoning is economic.
Advertising depends on a human reaching the page. A training crawler can consume the content without producing that visit, while an agent may fetch the page on a user’s behalf without displaying the publisher’s ad experience.
Search, by contrast, is treated as a discovery mechanism that can still send human traffic.
The three crawler categories represent three different economic relationships
Cloudflare’s Search, Training and Agent taxonomy is useful beyond its own dashboard.
A search crawler typically indexes content so users can discover and click it. A training crawler consumes content to improve a model. An agent retrieves a page while performing a task for a user.
Those behaviors can create very different value exchanges for a publisher.
Grouping them all under “AI bots” hides the commercial differences.
The new controls acknowledge that a publisher may rationally allow one behavior and reject another.
Robots.txt still expresses a preference, not a universal enforcement mechanism
Cloudflare is explicit about the limits of robots.txt.
A directive can communicate what a site owner wants, but it cannot identify the true operator behind every request or force a crawler that ignores the standard to comply.
Cloudflare’s network position lets it add another layer by identifying bot traffic and blocking training crawlers that do not qualify for the cooperative mixed-use treatment.
That combination of published preference and edge enforcement is what makes the new control more than a robots.txt generator.
Publishers should audit their current Cloudflare settings now
The immediate technical task is to verify what the site is actually configured to do.
If the objective is traditional search discoverability without AI-training permission, Search should remain allowed and Training should use Disallow AI Training rather than Block.
Teams should then inspect the resulting robots.txt and confirm that important search crawlers can still fetch representative URLs.
For Bing, the team must separately decide whether the NOARCHIVE tradeoff is acceptable until Microsoft adds robots.txt support.
Do not use Google-Extended as an AI Overviews removal switch
This is likely to be one of the most common implementation mistakes.
Google-Extended is a training and generative-AI-use control defined by Google; it is not a general “remove me from Google AI Search” directive.
Publishers that want to control appearance in AI Overviews, AI Mode or other generative Search features need to use the controls Google provides for those search experiences.
Cloudflare’s Disallow AI Training setting does not make that decision on the publisher’s behalf.
Do not select Block unless losing search crawling is intentional
The opposite mistake is potentially more damaging.
Cloudflare now says Block applies to mixed-use crawlers. A publisher selecting it because the label sounds like the strongest way to reject AI training can inadvertently stop Googlebot, Applebot or Bingbot from crawling for search.
That can undermine the discoverability the publisher was trying to preserve.
The semantic difference between Disallow and Block is therefore not cosmetic. It maps to fundamentally different network behavior.
The real innovation is purpose-specific crawler governance
For years, crawler policy was largely binary: allow the bot or block the bot.
AI has made that model inadequate because the same organizations can use web content for indexing, model training, generative summaries and user-directed agent activity.
Publishers increasingly want to consent to those purposes separately.
Cloudflare’s new framework attempts to translate that policy preference into technical controls while pressuring crawler operators to provide the transparency needed to make the distinction credible.
Google, Apple and Microsoft are moving toward a common principle, not a common implementation
All three companies have told Cloudflare that publishers should be able to reject training without sacrificing traditional search ranking.
The implementation details remain different.
Google uses Google-Extended. Apple uses Applebot-Extended. Microsoft currently uses NOARCHIVE and plans robots.txt support for early 2027.
Those differences matter operationally because the side effects are not identical.
A publisher needs to understand the platform-specific mechanism rather than assuming one Cloudflare toggle produces the same downstream result everywhere.
AI crawler policy is becoming part of technical SEO
This is no longer an issue that can be delegated entirely to legal or infrastructure teams.
The wrong crawler policy can affect search indexing, AI citations, model-training permissions and referral traffic at the same time.
Technical SEOs need to understand robots.txt directives, bot classification, CDN enforcement and the distinction between training and search-generation controls.
They also need to coordinate with editorial and business teams because the correct configuration depends on the publisher’s economic strategy.
The Bing gap shows why the standards work is unfinished
Cloudflare’s announcement is progress, but it also demonstrates how fragmented AI content controls remain.
A publisher should not need to learn one extended token for Google, another for Apple, a meta tag with additional visibility consequences for Microsoft and separate rules for every training-only crawler.
Cloudflare says it is continuing to work with operators and standards bodies, including the IETF, toward more interoperable controls.
Until those standards mature, infrastructure providers are effectively translating a publisher’s high-level preference into a growing set of vendor-specific mechanisms.
Publishers can finally separate Google Search from Google training—but not every AI tradeoff has disappeared
Cloudflare’s Disallow AI Training setting solves one of the most dangerous crawler-control problems for SEO teams: rejecting AI training no longer requires blocking Googlebot simply because Googlebot is a mixed-use crawler.
Google-Extended and Applebot-Extended provide dedicated preference channels, while Cloudflare can keep the underlying search crawlers available.
Bing remains the exception. Microsoft has committed to robots.txt support, but until early 2027 its current NOARCHIVE mechanism carries a real cost for publishers that value Bing Chat and Copilot citations.
The practical lesson is that “block AI” is no longer a sufficiently precise policy.
Publishers need to decide separately whether they want search indexing, model training, generative-search inclusion and agent access—and then configure each operator accordingly.
Cloudflare has made that separation substantially easier. The web’s crawler standards still have work to do.