Why the CMA’s publisher rules for Google may not go far enough
The UK’s Competition and Markets Authority (CMA) has opened its consultation on proposed conduct requirements for Google, aiming to give publishers more control and transparency over how their content is used in Google’s generative AI services. These are the first consultations of their kind under the UK’s digital markets competition regime, and they respond to a structural problem: publishers cannot realistically block Google’s crawler without losing search visibility, yet that same crawler now feeds AI features that return little traffic.
We support the CMA’s direction, but the proposal as written may not be enough to protect publishers or support fair competition in AI markets.
The SMS designation and what it enables
In January 2025, the Digital Markets, Competition and Consumers Act 2024 (DMCC) came into force, giving the CMA the ability to designate firms with substantial, entrenched market power as having Strategic Market Status (SMS). This designation allows the CMA to impose legally enforceable conduct requirements targeted at specific digital markets.
In October 2025, the CMA designated Google as having SMS in general search and search advertising, citing its 90 percent share of the UK search market. That designation explicitly covers AI Overviews and AI Mode, meaning the CMA can now impose requirements on how Google’s search ecosystem operates, with significant sanctions for non-compliance.
Why publishers can’t opt out
As the CMA itself states, publishers have no realistic option but to allow Google to crawl their content for general search, because of Google’s market power. The problem is that the same crawled content is also used for generative AI features that compete directly with publishers, often without attribution or compensation.
Blocking Googlebot is not a viable choice for publishers, because it would cut off the search traffic that supports ad-funded business models. That leaves publishers stuck: they must accept that their content will be used in AI Overviews and AI Mode, which return very little traffic in return.
This dynamic gives Google a structural advantage in the market for generative AI. Unlike other AI bot operators, Google can gather content for AI training and inference without meaningful risk of being blocked. Other AI companies are effectively disincentivized from negotiating fairly with publishers, because a dominant player can bypass compensation entirely.
Cloudflare data shows the scale of the imbalance
Data from Cloudflare’s network over a two-month period confirms the concern. Googlebot accessed significantly more unique pages than other major AI crawlers.
In rounded multiples, Googlebot saw:
- ~1.70x the unique URLs seen by ClaudeBot
- ~1.76x the unique URLs seen by GPTBot
- ~2.99x the unique URLs seen by Meta-ExternalAgent
- ~3.26x the unique URLs seen by Bingbot
- ~5.09x the unique URLs seen by Amazonbot
- ~14.87x the unique URLs seen by Applebot
- ~23.73x the unique URLs seen by Bytespider
- ~166.98x the unique URLs seen by PerplexityBot
- ~714.48x the unique URLs seen by CCBot
- ~1801.97x the unique URLs seen by archive.org_bot
Beyond raw crawl volume, the blocking patterns confirm the dilemma. Robots.txt files rarely disallow Googlebot in full, because of its role in driving search referrals. Partial disallows typically only affect areas irrelevant to SEO, such as login endpoints. When publishers use Web Application Firewall (WAF) rules or tools like Cloudflare’s AI Crawl Control—which integrates with the Application Security suite—they are nearly seven times more likely to block other AI crawlers like GPTBot and ClaudeBot than Googlebot or Bingbot.

The reluctance to block Googlebot is rational for publishers, but it gives Google an unfair advantage. It can gather data for AI functions with little fear of losing access, and it has minimal incentive to pay for content it already receives for free. This undermines the emergence of a well-functioning marketplace where AI developers negotiate fair value for publisher content.
What the CMA proposes
On January 28, 2026, the CMA published four sets of proposed conduct requirements for Google, including one specifically for publishers. The proposed rules address three concerns: publishers lack sufficient choice over how their content is used in AI-generated responses, lack transparency into that use, and do not receive effective attribution.
The requirements would mandate that Google provide publishers with “meaningful and effective” control over whether their content is used in AI features. Google would also be prohibited from taking actions that undermine those controls, such as intentionally downranking content in search results. In addition, Google would need to publish clear documentation on how it uses crawled content for generative AI and what each publisher control covers in practice. Finally, Google would have to ensure effective attribution and provide publishers with detailed, disaggregated engagement data, including impressions, clicks, and “click quality,” so they can evaluate the commercial value of allowing their content in AI-generated summaries.
We welcome the CMA’s recognition that publishers need better protections, and we agree that meaningful opt-out mechanisms are essential. But given the scale of the imbalance, the proposed requirements may not be sufficient to level the playing field. A system that separates search crawling from AI crawling is needed, so that publishers are not forced to choose between visibility in search and control over AI use of their content. Without that separation, the competitive advantage Google holds will persist even under the new rules.
Proposed remedies leave publishers dependent on Google’s terms
Cloudflare supports the CMA’s goal of giving publishers better options, but argues the current requirements fail to address the core problem: publishers are still forced to rely on Google’s proprietary opt-out mechanisms, which are tied to Google’s platform and defined by Google’s conditions. That setup gives publishers no genuine, autonomous control. When one company dictates the rules, builds the technical controls, and decides the scope of application, content creators do not gain effective control; they simply remain dependent. This arrangement also undermines competitive innovation.
The proposed framework narrows publisher choice. If Google introduces new opt-out controls, publishers could not use external tools to block Googlebot without risking their Search visibility. Under the CMA’s current proposal, content creators would still have to let Googlebot scrape their sites, with no enforcement mechanisms to fall back on and little visibility if Google ignores stated preferences. Properly enforcing these requirements, Cloudflare argues, would be extremely onerous for the CMA — and there is no guarantee publishers would trust the outcome.
Cloudflare says its customers report that Google’s current opt-out tools, including Google-Extended and nosnippet, have not prevented content from being used in ways publishers cannot control. These tools also offer no route to fair compensation.
Cloudflare’s responsible AI bot principles hold that all AI bots should have a single, declared purpose so site owners can make clear, informed decisions about who gets access and why. Competitors like OpenAI and Anthropic follow this model; Google does not, the company notes. Googlebot serves multiple purposes — search indexing, AI training, and inference/grounding — leaving publishers unable to make granular choices. A new opt-out mechanism alone would not fix that.
The only meaningful solution, in Cloudflare’s view, is splitting Googlebot into separate crawlers by purpose. That would let publishers allow crawling for search indexing — essential for driving traffic — while blocking access for generative AI training and related features.
Why separating Googlebot is both feasible and necessary
The CMA acknowledged the separate-crawler remedy was an “equally effective intervention” but rejected it after Google said it would be too burdensome. Cloudflare disagrees. Google already operates nearly 20 other crawlers for distinct functions, so adding purpose-specific ones for indexing and AI usage is technically straightforward.
A separation remedy would not add traffic load; site owners might even see reduced load if they choose to block AI crawling. With separate crawlers, publishers could block unwanted AI access before Google fetches content, instead of relying on post-hoc, Google-managed controls. This also gives publishers the ability to negotiate the conditions under which their content gets used.
Beyond publisher benefits, Cloudflare argues the remedy levels the playing field between Google and other AI companies. It notes support from Daily Mail Group, the Guardian, and the News Media Association. Mandatory crawler separation is not a penalty for Google or a drag on AI investment; in fact, it prevents Google from using its search dominance to tip the AI market in its favor. Decoupling indexing from AI training keeps AI development anchored in fair-market competition, not the exploitation of a single hyperscaler’s position.
Cloudflare calls on the CMA to press forward with rules that treat Google like any other AI developer when it comes to content access. That approach would give publishers real agency over use and monetization of their work. Cloudflare says it will participate in consultations and offer evidence that helps shape conduct requirements that are targeted, proportional, and effective — and that keep the Internet a fair marketplace for content creators and smaller AI players, rather than a few large platforms.



