AI crawler activity gets new Radar filters for purpose and industry
Search engines once operated on a straightforward exchange: crawl a site, index it, and send traffic back in return for appearing in relevant results. AI platforms have broken that compact. Users asking a chatbot or an AI-augmented search engine often get an answer without ever clicking through to the source — when a source link is provided at all. For publishers, that means fewer page views and less ad revenue from queries that used to arrive via search.
Cloudflare Radar has been tracking this shift since July 1, when it launched crawl/refer ratios. Those figures compare the volume of HTML crawling requests from each platform's crawler against the human traffic those platforms refer back to sites. Now, Radar is adding two more lenses to its AI Insights page: a breakdown of AI bot traffic by crawl purpose and an industry filter for comparing activity across verticals.
Crawl traffic now has a purpose
Since LLMs entered the mainstream in late 2022, the dominant reason AI platforms crawl the web has been to scrape content for model training. That activity can be aggressive and sometimes ignores directives in robots.txt files. But training is no longer the only motive. AI platforms are also crawling to build search indexes that compete with classic engines, and some crawl pages in direct response to a user prompt — for instance, checking flight options for a vacation query.
The new Crawl purpose selector on the AI bot & crawler traffic card distinguishes four categories: Training, Search, User action, and Undeclared for crawlers whose operators have not disclosed an intended use.
Selecting a purpose updates the HTTP traffic by bot graph to show the top five most active crawlers for that category over the selected period. Choosing User action, for example, shows OpenAI's ChatGPT-User bot accounting for roughly three-quarters of requests across the first 28 days of July 2025. The traffic shows a clear daily cycle, implying regular use of ChatGPT to answer live queries, with activity gradually climbing through the month. If ChatGPT-User is stripped out, Perplexity-User follows a similar daily pattern.
A new Crawl purpose graph displays traffic trends across all purposes. Training traffic dominates at nearly 80% of all AI bot crawling and is erratic, with no obvious cyclical rhythm. User action and Undeclared traffic show visible daily cycles but together account for under 5% of AI bot requests in the same late-July window.

The Data Explorer view for the AI Bots & Crawlers dataset now accepts a Crawl purpose breakdown, so activity can be compared over time. The same view supports grouping by User agent and filtering by Crawl purpose, which exposes trends across a wider set of bots than the top-five shown on the main page. Time-period comparisons are also available in Data Explorer.
Industry views put your numbers in context
Site operators can watch their own logs for aggressive scraping and can see how often AI platforms send traffic back to their domains. What is harder to gauge is whether that activity is out of line with peers. New industry set filters on the HTTP traffic by bot graph and the Crawl-to-refer ratio table in AI Insights make that comparison possible.
A drop-down at the top of the AI bot & crawler traffic card lets you pick an industry set. The choice updates the HTTP traffic by bot and Crawl purpose graphs along with the Crawl-to-refer ratio table. Selecting a crawl purpose further narrows the bot traffic graph.

The resulting patterns differ sharply by industry. With no industry or purpose selected across the first week of August, ClaudeBot and GPTBot together account for nearly half of crawling activity, and among the top five only Meta-ExternalAgent shows anything close to a regular pattern. In the unfiltered view, Anthropic's crawl-to-refer ratio is roughly 50,000:1, with OpenAI at 887:1 and Perplexity at 118:1.
Filtering to the News and Publications industry set spreads traffic more evenly across the top five — ChatGPT-User at 14.9% and GPTBot at 17.4% — and the presence of ChatGPT-User suggests readers are querying AI assistants about current events. Crawl-to-refer ratios are markedly lower for these sites: Anthropic at about 2,500:1, OpenAI at 152:1, and Perplexity at 32.7:1.

For the Computer and Electronics industry set, GPTBot is again the most active crawler but Amazonbot moves into second place; together those two represent over 40% of crawl traffic. ClaudeBot and Meta-ExternalAgent each hold a 13.9% share, with ByteSpider completing the top five. Ratios here sit between the unfiltered and news numbers: Anthropic at about 8,800:1, OpenAI at 401.7:1, and Perplexity at 88:1.

Data Explorer also supports breakouts by vertical and industry, where a vertical is a pre-defined group of related industries. Both the Crawl purpose and User agent breakdowns can be filtered this way. For example, sites in the Cryptocurrency industry under the Finance vertical see crawling from a number of bots, but three-quarters of that traffic in the first week of August came from just four crawlers, with roughly 80% of requests aimed at gathering training data.
The industry sets shown on the main AI Insights page are manually curated collections of related industries. Clicking through to Data Explorer from one of those views pre-populates the industry selector with the constituent entries — for instance, the Gaming & Gambling set expands into its component industries.

What comes next
AI crawler traffic is now a permanent part of the web, and the range of reasons those crawlers visit keeps growing beyond LLM training. Work on standards such as content signaling aims to give publishers a way to declare how automated agents may use their content, but those proposals still need to be standardized and adopted on both sides. In the meantime, Radar's AI Insights continue to expand as the crawler landscape evolves.



