A Fresh Look at Crawler Waste
Most web traffic is not human. Cloudflare's position on the network — sitting in front of more than one in six websites — gives it a continuous view of the flow between users, origins, and automated systems. Nearly half of all global traffic comes from bots. Much of that is malicious, and stopping it before it reaches customer origins prevents a cascade of unnecessary database queries and dynamic page generation on infrastructure that is far less efficient than Cloudflare's own edge.
But the good bots are a problem too. Search engine crawlers and other indexing services account for more than 5% of global Internet traffic, and they are essential to making the web navigable. The issue is that much of their work is redundant. After a year of observing crawler behavior, Cloudflare estimates that 53% of good bot traffic is wasted — revisits to pages that have not changed since the crawler last saw them.
The math on the environmental opportunity is striking. The Boston Consulting Group estimates the Internet generates about 1 billion metric tonnes of carbon per year, roughly 2% of global output. If good bots are 5% of traffic and 53% of that is excessive crawling, eliminating the waste could save on the order of 26 million tonnes of carbon annually. For context, the U.S. Environmental Protection Agency's equivalency calculator puts that at the impact of 31 million acres of forest, six coal-fired power plants, or 5.5 million passenger vehicles.
Knowing When to Crawl
Search engines have gotten very good at running efficient data centers and servers. The remaining inefficiency is in the crawl strategy itself. Crawlers cannot know when a page changed without checking, and the existing mechanisms for telling them are clunky. Site owners can request a recrawl from Google, for example, but the visit happens "a few days to a few weeks" later. Coordinating recrawl requests across multiple search engines means tracking when each one last visited — and even then, the request model lacks explicit change data.
Sitemaps are the better foundation. The protocol is open, well-defined, and supports large sites with millions of URLs. But building and maintaining accurate sitemaps is technically complex, especially for sites built on disparate technologies. Many sites simply cannot keep one current enough to be useful.
Cloudflare's position changes that calculation. From the edge, it sees every page it serves and knows which ones changed — by hash or timestamp — and when a crawler last visited. That visibility lets it construct a per-crawler, automatically updated list of URLs that have changed since that crawler's previous visit. Each search engine gets exactly what it needs, nothing more, and the origin sees zero additional load. The sitemap protocol's priority field can also be informed by real visitor traffic, giving crawlers a hint about which pages are significant enough to index promptly.
Crawler Hints
Cloudflare is announcing Crawler Hints, a mechanism that supplies search engine crawlers with high-quality signals about content changes on Cloudflare-backed sites. The intent is to let crawlers time their visits precisely, skip unchanged pages, and reduce resource consumption across crawler infrastructure, customer origins, and Cloudflare's own network. A side benefit: fresher crawler data should mean more relevant search and indexing results.
The protocol itself is open and not tied to Cloudflare. The company hopes other hosting providers and similar services will adopt it, and it plans to work with the search and hosting communities on refinements. Operational details remain, including how a crawler identifies itself to retrieve its personalized list, but the direction is set: stop guessing about content freshness and let the edge tell you.
Who Benefits
For site owners, Crawler Hints should reduce bot traffic hitting the origin, which means less resource consumption and a smaller carbon footprint. It should also eliminate competition between bots and human users for origin resources, improving performance. Fresher content reaching search indexes can influence rankings and user experience.
For end users, bot-fed services — search engines, pricing tools, and similar features that most people use daily without thinking about it — will return more current results because their underlying data refreshes when the source actually changes.
For the Internet at large, the energy savings are the headline. Cutting needless crawls reduces computing work across the entire chain: the crawler's fleet, the edge, and the origin. The model is not perfectly simple — real-world efficiency gains will not map one-to-one to carbon reductions — but the scale of the opportunity is real.
Cloudflare is already lining up partners. Yandex, DuckDuckGo, and the Internet Archive's Wayback Machine have all expressed support for the initiative. The Internet Archive, which previously partnered with Cloudflare on the "Always Online" service, expects Crawler Hints to let it focus server and bandwidth resources on pages that have actually changed rather than re-scanning the rest of the web.
Yandex prioritizes long-term sustainability over short-lived success, and joins the global community in its pursuit of climate change mitigation. As a part of its commitment to quality service and user experience, Yandex focuses on ensuring relevance and usability of search results. We believe that a Cloudflare's solution will strengthen search performance by improving the accuracy of returned results, and look forward to partnering with Cloudflare on boosting the efficiency of valuable bots across the Internet.
"DuckDuckGo is supportive of anything that makes search more environmentally friendly and better for end users without harming privacy. We're looking forward to working with Cloudflare on this proposal."– Gabriel Weinberg, CEO and Founder, DuckDuckGo.
Nearly a year ago (the Internet Archive’s Wayback Machine partnered with Cloudflare) to help power their "Always Online" service and, in turn, to have the Internet Archive learn about high-quality Web URLs to archive. That win-win partnership has been a huge success for the Wayback Machine and, in turn, our partners, as it has helped ensure we better fulfill our mission to help make the Web more useful and reliable by backing up, and making available for future generations, much of the public Web. Building on that ongoing relationship with Cloudflare, the Internet Archive is thrilled to start using this new "Crawler Hints" service. With it, we expect to be able to do more with less. To be able to focus our server and bandwidth resources on more of the Web pages that have changed, and less on those that have not. We expect this will have a material impact on our work. The fact the service also promises to reduce the carbon impact of the Web overall makes it especially worthwhile and, as such, we are proud to be part of the effort.-- Mark Graham, Director, the Wayback Machine at the Internet Archive



