Content owners who want search traffic but not AI training have been stuck with a binary choice, because the largest crawlers on the web are mixed-use: one bot serving both a search index and model training. Blocking the bot blocks the discovery along with the training.
Cloudflare's new Disallow AI Training setting decouples the two. It keeps a site indexed for search while refusing training on the same crawler, and Apple, Google, and Microsoft honor it or have committed to honoring it within a specified time frame.
Summaries are the next front. A site-wide yes-or-no is too coarse for AI summaries, where how much of a page surfaces matters as much as whether it surfaces at all. An opt-out for summaries is already among the requirements Cloudflare has set for mixed-use crawler operators, and by early next year the goal is a single Cloudflare-side control over how much content is included, rather than a separate arrangement with each operator.
Why robots.txt alone can't do this
The demand for search is near-universal: fewer than 1% of Cloudflare sites block search bots. Training draws the opposite response, with 17% of sites enabling some mechanism to block it. Those figures are the argument for granular controls instead of a blanket "Block AI."
A robots.txt directive can express a preference, but it cannot identify who is crawling, determine why, or stop a crawler that disregards it. A network can: publish the preference, identify the crawler, classify its purpose, block the ones that ignore the directive, and report on operator behavior via Radar.
Blocking, though, only removes a crawler — it does not change crawler behavior. The preferred end state is operators that don't force the choice at all, which is why Cloudflare has been in direct talks with them since July. The responses were encouraging: nearly all agreed site owners should have control and transparency over how their content is used, and that their choices should be respected.
The Accountable designation
Accountable is a Cloudflare designation that recognizes capabilities shipping today alongside concrete, time-bound commitments. Operators must meet or commit to meeting four requirements:
- An opt-out from AI training for site owners, via robots.txt or a similar standard.
- An opt-out from AI summaries, set with the operator directly and, next year, through Cloudflare.
- URL-level visibility into which pages were made available for training, plus metrics on how content appeared in search.
- Assurance that opting out of AI training will not affect traditional search results.
Apple, Google, and Microsoft all qualify. Each pairs existing capabilities with commitments to deliver what is still in development.
How the controls are classified
Cloudflare sorts bots by behavior rather than identity, and one bot can carry more than one behavior. Three are exposed as controls:
- Search — crawling to build a search index.
- Training — crawling to train or fine-tune a model.
- Agent — user-directed agents fetching a page on a human's behalf, such as chat fetch bots and browser-use agents.
A crawler that does Search and Training at once is what created the tradeoff. Disallow AI Training, named for the Disallow: directive it writes into robots.txt, exists so that Accountable mixed-use crawlers need not be blocked outright — they stay allowed for search.
Existing settings change meaning. Block and "Block on pages with ads" previously skipped mixed-use crawlers precisely because blocking them endangered discoverability. With Disallow AI Training available, both now cover all training crawlers, mixed-use included. The four Training settings, applied at the domain level, are:
- Allow: everything is permitted unless another setting or a WAF rule says otherwise.
- Disallow AI Training: Bot Preference Sync publishes the no-training preference in robots.txt. Accountable mixed-use crawlers remain allowed for search; every other training crawler is blocked, including the training-only crawlers from Amazon, Anthropic, Meta, and OpenAI — none of which affects search. Available only for Training, not Search or Agent.
- Block on pages with ads: crawlers, mixed-use included, are blocked only on pages detected serving an ad.
- Block: all crawlers, mixed-use included, are blocked.
There is deliberately no ads-only variant of Disallow AI Training. The setting works by publishing a robots.txt preference, and an ads-only preference cannot be expressed that way: Cloudflare can detect which pages serve ads, but the list is too large and too volatile to enumerate in robots.txt.
Agents get no Disallow setting for now. They do not create the same discoverability tradeoff as mixed-use crawlers, and no well-established directive exists yet for refusing them. Cloudflare says it will revisit this as standards such as ai-prefs mature.
What happens on September 15
The Bot Management and AI Crawl Control changes are:
- Block and Block on pages with ads now apply to mixed-use crawlers — Applebot, Bingbot, Googlebot — so either setting hits search alongside training. Use Disallow AI Training to stop training while keeping search.
- "Block AI Bots" is deprecated in favor of the Search, Training, and Agent controls.
- Managed Robots.txt is deprecated in favor of Bot Preference Sync, with existing users migrated.
- Disallow AI Training joins the recommended configuration for certain new domains.
- Existing customers' preferences are migrated to the new controls.
In almost every case there is nothing to do; current settings carry over automatically. Site owners who want mixed-use crawlers gone entirely must now say so explicitly by selecting Block, which stops Applebot, Bingbot, and Googlebot — search included.
Domains that never used the Search/Training/Agent controls migrate according to their legacy Block AI Bots setting:
| (Legacy) “Block AI” setting |
(New) Search setting |
(New) Training setting |
(New) Agent setting |
|---|---|---|---|
| Disabled (unselected) | Allow | Allow | Allow |
| Block | Allow | Disallow AI Training | Block on pages with ads |
| Block on pages with ads | Allow | Disallow AI Training | Block on pages with ads |
Domains that had already configured the granular controls keep the practical effect of their selections under the new definitions, with prior Block or Block on pages with ads choices for Training becoming Disallow AI Training.
| Control | Legacy setting | New setting |
|---|---|---|
| Search | Allow | Allow |
| Block | Block | |
| Block on pages with ads | Block on pages with ads | |
| Training | Allow | Allow |
| Block | Disallow AI Training | |
| Block on pages with ads | Disallow AI Training | |
| Agent | Allow | Allow |
| Block | Block | |
| Block on pages with ads | Block on pages with ads |
From September 15, onboarding a new domain presents one of two presets, split by whether the site earns advertising revenue. Ad revenue requires a human to actually see the page; a training answer replaces that visit, and an agent fetches the page with no one there to view the ads. Presets for ad-supported sites are therefore more restrictive. Either can be changed during onboarding or at any later point.
| Setting | Site does not monetize using ads | Site is monetized using ads |
|---|---|---|
| Preference Sync | Enabled | Enabled |
| Search | Allow | Allow |
| Training | Allow | Disallow AI Training |
| Agent | Allow | Block on pages with ads |
Recommended settings for new domains

Crawler-by-crawler behavior under Disallow AI Training
The mixed-use crawlers fall into two groups. Applebot, Bingbot, and Googlebot are all classed as Accountable, and their operators — Apple, Google, and Microsoft — have committed to the same publisher-choice and transparency principles. With Disallow AI Training selected, these crawlers keep indexing the site for search; switching to Block stops them completely.
The relevant crawlers from Amazon, Anthropic, Meta, and OpenAI are also categorized as Accountable. Because those organizations run separate Search and Training crawlers, Cloudflare can block the Training crawler without touching search.
Applebot
A robots.txt Disallow rule for Applebot-Extended is the training opt-out. Preferences about AI Summaries can be expressed today through the nosnippet directive in page HTML, and content can be marked as paywalled so it is excluded from generative output. Applebot has no URL-level inspection tool yet; Apple has shared details of an in-progress solution for next year with Cloudflare. Apple states that disallowing training does not affect search ranking.
Googlebot
Google’s training opt-out is a robots.txt Disallow for Google-Extended, plus a toggle in the webmaster portal that removes a site’s content from generative search results. Googlebot also reports metrics on search results and AI summary results. Google has described additional URL-level transparency tools for site owners related to Google-Extended, expected in the coming weeks, and states that disallowing Google-Extended has no effect on a site’s inclusion or ranking in Google Search.
Bingbot
Bingbot’s granular controls sit in Webmaster Tools. The current way to express a training preference is the NOARCHIVE meta tag, which keeps content out of Bing Chat answers and out of training for Microsoft’s generative AI foundation models. Microsoft is extending this and building robots.txt support for a domain/site-level “no training” preference, targeted for early 2027. Cloudflare customers opting out today can combine NOARCHIVE with the Block URLs or Content Removal tool. Microsoft states that NOARCHIVE does not affect search ranking.
Until that robots.txt support ships, selecting Disallow AI Training will not automatically carry a no-training preference to Bing. That matches the previous Training Block setting, which did not apply to mixed-use crawlers such as Bingbot.
Accountability tracking and standards work
Cloudflare says it will keep engaging with AI crawler operators as these controls evolve, and Radar publicly tracks the controls, transparency, and reporting offered by Accountable operators.
The stated goal is dual agency: crawlers get access to the open web, and the people who create it get meaningful control over how their work is used. Reaching that balance is framed as a shared effort involving infrastructure providers, content creators, technology companies, and standards bodies such as the IETF, which must turn these principles into open, interoperable standards.
Why summaries are a different problem from training
Training and AI Summaries raise separate questions. Training governs whether content can be used to build models; summaries govern how people discover, evaluate, and visit a business.
Summary opt-outs are the first step, and the Accountable operators either offer that capability or are finishing work to provide it — the baseline being that a site owner can say no. But a site-wide allow/prohibit switch for summaries is still blunt: the right call depends on the site, the content, and the business outcome.
For publishers, training raises questions of control, compensation, and the sustainability of original content, while summaries pose a more immediate distribution question — whether someone visits the publisher’s site or consumes the answer inside a search or AI experience. For other businesses, summaries often sit between a prospective customer and the website, answering a question, comparing alternatives, or recommending a product before any visit happens.
Reported data shows mixed effects. More than half of consumers read summaries in Search, and those who do are over 40% more likely to end their search after reading one, which can cut the visits a site receives. At the same time, consumers referred by AI Search convert at between three times and over five times the rate of those referred by traditional search — fewer visits, but with much greater intent.
Neither outcome is inherently good or bad. An ad-funded publisher may optimize for audience volume; a retailer may prefer fewer, more likely buyers. Cloudflare’s position is that its role is to supply the visibility and control for site owners to decide, not to decide for them. Its next focus is helping them understand how summaries affect their businesses and giving them more control over how much content can be used, with open standards such as ai-prefs as an important part of that.
Feedback on this work can be sent to [email protected]. The new controls are available on all plans and can be configured at the domain (zone) Security Settings.
Progress requires infrastructure providers, content creators, technology companies, and standards bodies such as the Internet Engineering Task Force (IETF) working together to translate these principles into open, interoperable standards.



