Who Is Crawling Your Site in 2025?
Web crawlers have been part of the internet’s backbone since the early days of search. The World Wide Web Wanderer appeared in 1993, followed by search engines like JumpStation and WebCrawler that used crawlers to index content. Their job was straightforward: scan websites so pages could appear in search results. That role has expanded dramatically with the rise of AI, which now relies on the same mechanisms to harvest training data.
Cloudflare Radar data shows that bots account for roughly 30% of global web traffic, sometimes exceeding human traffic in certain regions. These bots range from benign search indexers like Googlebot to malicious actors engaged in credential stuffing or unauthorized scraping. A newer category—AI crawlers—has emerged to collect data for training large language models and other AI systems, raising questions about content rights and infrastructure load.
AI Crawlers: A Shifting Landscape
Comparing May 2024 to May 2025, the AI crawler field has been reordered. GPTBot from OpenAI surged from a 5% share to 30%, while Meta-ExternalAgent made a strong debut at 19%. This growth came at the expense of former leader Bytespider, which dropped from 42% to 7%, and other crawlers like ClaudeBot and Amazonbot also lost ground.


Key crawlers in this space include:
- GPTBot – OpenAI’s crawler for training and improving models like ChatGPT.
- ClaudeBot – Anthropic’s crawler for the Claude AI assistant.
- Meta-ExternalAgent – Meta’s bot, likely collecting data for LLM training or fine-tuning.
- Amazonbot – Amazon’s crawler for search and AI applications.
- Bytespider – ByteDance’s AI data collector, often tied to Ernie or TikTok-related AI.
- Applebot – Apple’s crawler for Siri and Spotlight, with possible AI uses.
- OAI-SearchBot – OpenAI’s search-focused crawler for real-time web info.
- ChatGPT-User – Represents API-based or browser usage of ChatGPT.
- PerplexityBot – Crawler powering Perplexity.ai’s real-time answer engine.
Website owners can signal preferences to these bots via robots.txt, though compliance is voluntary. Cloudflare has introduced tools like AI Audit to help creators enforce their policies when crawlers ignore them.
Overall Crawling Growth
A broader analysis covering both search and AI crawlers shows a clear upward trend. Using a fixed set of customers to eliminate growth bias, the data reveals an 18% increase in crawling traffic from May 2024 to May 2025. The figure jumps to 48% if new Cloudflare customers added during that period are included. Peak activity occurred in April 2025, with a 32% increase compared to May 2024.

The traffic patterns follow seasonal human internet behavior. Crawling activity dipped during the Northern Hemisphere summer of 2024, with August and September as the least active months, then rose in November alongside increased consumer activity.
Googlebot Dominance and Growth
Googlebot remained the top crawler throughout the entire period, with traffic up 96% from May 2024 to May 2025. Crawling peaked in April 2025 at 145% higher than May 2024 levels, coinciding with Google’s rollout of AI Overviews in search results, which began in the US in May 2024 and later expanded to other countries.

Daily data for Google’s crawlers shows two notable dips. The first occurred on December 14, 2024, around a Google Search update. The second, from May 20 to May 28, 2025, coincided with the US rollout of AI Mode on Google Search, though the timing may be coincidental. The bulk of Google’s crawling activity comes from Googlebot and GoogleOther, the latter introduced in 2023 for research and development purposes.

The Crawler Leaderboard Shakes Up
Looking at raw request share among a cohort of more than 30 AI and search crawlers observed by Cloudflare in May 2024 versus May 2025, the ranking has shifted considerably. GPTBot climbed from the #9 spot to #3, while ClaudeBot and Bytespider lost significant ground. The most notable mover overall, however, was Googlebot, which expanded its share from 30% to 50% of all crawling traffic in this group.
| Rank | Bot name | Share May 2024 | Share May 2025 | Δ percentage-point change | Raw requests growth (May 2024 to May 2025) |
|---|---|---|---|---|---|
| 1 | Googlebot | 30% | 50% | +20 pp | 96% |
| 2 | Bingbot | 10% | 8.7% | -1.3 pp | 2% |
| 3 | GPTBot | 2.2% | 7.7% | +5.5 pp | 305% |
| 4 | ClaudeBot | 11.7% | 5.4% | -6.3 pp | -46% |
| 5 | GoogleOther | 4.4% | 4.3% | -0.1 pp | 14% |
| 6 | Amazonbot | 7.6% | 4.2% | -3.4 pp | -35% |
| 7 | Googlebot-Image | 4.5% | 3.3% | -1.2 pp | -13% |
| 8 | Bytespider | 22.8% | 2.9% | -19.8 pp | -85% |
| 9 | Yandex | 2.8% | 2.2% | -0.7 pp | -10% |
| 10 | ChatGPT-User | 0.1% | 1.3% | +1.2 pp | 2,825% |
| 11 | Applebot | 1.9% | 1.2% | -0.7 pp | -26% |
| 12 | Timpibot | 0.3% | 0.6% | +0.3 pp | 133% |
| 13 | Baiduspider | 0.5% | 0.4% | -0.1 pp | 7% |
| 14 | PerplexityBot | <0.01% | 0.2% | +0.2 pp | 157,490% |
| 15 | DuckDuckBot | 0.2% | 0.1% | -0.1 pp | -16% |
| 16 | SeznamBot | 0.1% | 0.1% | 2% | |
| 17 | Yeti | 0.1% | 0.1% | 47% | |
| 18 | coccocbot | 0.1% | 0.1% | -3% | |
| 19 | Sogou | 0.1% | 0.1% | -22% | |
| 20 | Yahoo! Slurp | 0.1% | 0.0% | -0.1 pp | -8% |
The data reveals two dominant trends over the 12-month period.
AI Crawlers: Steep Climbers and Fast Fallers
GPTBot (OpenAI) saw its share jump from 2.2% to 7.7% (a +5.5 percentage point change) on the back of a 305% rise in raw requests. This surge reflects the growing appetite for web content to train large language models. A sibling crawler, ChatGPT-User, saw requests balloon by 2,825%, capturing a 1.3% share as users and API clients increasingly pull live web content. PerplexityBot posted the highest growth rate of any crawler in the cohort—a staggering 157,490% increase in raw requests—though it still holds a relatively small 0.2% share.
Not every AI crawler gained traction. Anthropic’s ClaudeBot dropped from 11.7% to 5.4% of traffic and saw a 46% decline in requests. Bytespider fell 85% in request volume, sliding from #2 to #8 in the share rankings (now at 2.9%). Both Amazonbot and Applebot also posted declines in share and raw requests (–35% and –26%, respectively).
Google Tightens Its Grip
Googlebot’s rise to a full half of all crawler traffic points to more than just traditional search indexing; it likely also supports AI-driven features like AI Overviews. The GoogleOther crawler (introduced in 2023) also increased its crawling traffic by 14%, and Googlebot-News grew a notable 71% in requests even though it didn't crack the top 20. This expansion comes as Google invests heavily in blending AI with its core search product.
Among legacy search engines, Microsoft’s Bingbot saw its share dip from 10% to 8.7% (a –1.3 pp change), though its raw request total still eked out a 2% gain.
One bot notably absent from the current top 20 is FriendlyCrawler, which ranked #14 in May 2024 with a 0.2% share but saw a 100% drop in requests and now sits at #35. Its owner and purpose remain unclear, though it has been observed indexing website content.
robots.txt Preferences: Blocking vs. Allowing
Analysis of Cloudflare Radar data from June 6, 2025, covering 3,816 domains from the top 10,000 list, found that 546 sites (roughly 14%) had explicit “allow” or “disallow” directives (fully or partially) aimed at AI bots specifically.

Disallow rules are far more common than allow rules. GPTBot leads both categories: it’s blocked by 312 domains (250 fully, 62 partially) and also explicitly allowed by 61 domains (18 fully, 43 partially). Other frequently blocked bots include CCBot and Google-Extended. The fact that so few sites grant open access—and when they do, it’s usually for limited sections—suggests widespread caution. Bots not mentioned in a site’s robots.txt are allowed by default, which leaves many domains in a gray area regarding newer or less transparent crawlers.
There is a growing recognition that robots.txt isn't always a reliable control for AI crawlers. Some site owners don't think to use it for this purpose, and others question whether newer bots even respect the rules. That’s part of why the ecosystem is moving toward more enforceable controls, such as Web Application Firewalls, alongside passive signals.
Note: Comparing crawler traffic involves matching user-agent tokens in robots.txt files against user-agent strings in HTTP requests. Not all tokens map directly; Google-Extended, for example, signals a purpose—whether content can be used for AI training—rather than identifying a distinct user agent. That means traffic from standard Google user-agents like Googlebot still appears in request logs for sites using such tokens, as described in RFC 9309.
As AI crawlers continue to reshape the web, the shift from search indexing toward AI model training is clear in the traffic data. The rise of Google and OpenAI bots, combined with the move toward stronger, enforceable blocking methods, signals how websites are adapting their access controls for an AI-driven future.



