User Insights, introduced last month, was built to answer a simple question: what are people actually doing with AI? The dashboard shows which users, applications, tasks, and models drive traffic, and flags anomalies in user and agent behaviour that point to unexpected or runaway spending.

The latest update adds the missing half of that picture. Model names and request counts say where traffic goes, not what work sits behind it — a code review, a research task, or an agent chaining calls to finish a job. Identical token counts can represent very different workloads, and model choice cannot be judged without knowing the task. User Insights now surfaces conversations where the selected model appears more capable than the task needs, shows which users and agents drive that usage, and relates task, model, cost, and conversation patterns. The capabilities are free for AI Gateway users.

Why usage is hard to interpret

Take a team that routes internal AI traffic through AI Gateway. Weeks later, spend is climbing and some requests feel slower than expected.

Several explanations are possible: developers taking on more complex coding work, agents issuing too many follow-up calls to complete a task, or a small set of users and agents accounting for a disproportionate share of usage. Tokens and request counts alone cannot distinguish between these patterns. Teams need to know what the traffic represents before changing a model, workflow, or routing rule.

The model overkill view

The overkill view identifies conversations where the chosen model is more capable than the work requires — for instance, simple formatting or summarization requests sent to a high-capability reasoning model. From there, teams can see which users, agents, or applications exhibit the pattern and investigate the underlying tasks. Common causes include a model that is simply the default, users unsure which model to pick, or an agent configured to use one model for every step.

BLOG-3566 2.png
Model fit overview

The view is not a leaderboard and never recommends a replacement model on its own. It is a prompt for better questions: whether the model fits the task, whether extra capability improves the result, whether a faster or cheaper model would produce an equivalent outcome, and whether the issue is confined to one workflow, user, or agent. Cost, latency, token usage, and conversation turns provide the evidence for a decision.

These signals feed two adjacent features. Potential Savings highlights requests that a faster or less expensive model could likely handle without hurting output quality, while the Auto Router applies task and model-fit signals automatically, removing the need for a separate routing rule per workload. The Auto Router is launching in public beta with this release.

BLOG-3566 3.png
Potential Savings view

Model fit is a comparison, not a rule. A difficult coding or research task may justify a capable reasoning model; a short summary or simple classification may not. The aim is not to push every request to the cheapest option, but to check whether the selected model suits the work.

BLOG-3566-replace-image.png
Users with Overkill Spend

Task categories and conversation turns

Task analysis groups conversations by the kind of work they represent, with initial categories covering coding, research, writing, summarization, and data analysis. This is context a list of model names cannot provide: one engineering team may skew toward coding and debugging, another toward research and summarization, and either may find that a surprising volume of traffic is simple work sent to a high-capability model. Because the data comes from traffic already passing through AI Gateway, teams can test whether a model is being used for work it suits or whether a default is being applied too broadly.

BLOG-3566 5.png
Categorical breakdown of tasks within a fictitious organization

Turns analysis covers how much back-and-forth a task takes. Some work finishes in one exchange; other work needs several rounds of questions, corrections, and follow-ups. A long conversation is not inherently bad, particularly for complex tasks, but a simple task that keeps consuming multiple turns is a signal to examine the prompt, the model, or the workflow. The first request is only part of the cost — the time, tokens, and money spent before completion matter too, and comparing those figures exposes workflows that run longer or cost more than expected.

BLOG-3566 6.png
Turn analysis on sessions

From insight to automatic routing

Once an overkill pattern is confirmed across task, cost, latency, and turn data, it becomes a candidate for automatic routing. A team might see that most of its traffic is summarization and formatting, that those requests go to a large reasoning model, and that most conversations finish in a single turn — a concrete workload to evaluate.

The Auto Router, now in closed beta, uses conversation trajectory, task category, task complexity, and model-fit signals to route requests automatically while accounting for cost. Rather than defining a rule per workload, beta customers let it choose among the models available to their application. It does not funnel everything to the cheapest model; complex coding or research work can still land on a more capable one, while simpler tasks may go to a faster or less expensive option. The Auto Router is built on the same task and conversation signals that power User Insights.

How classification works

Each conversation receives an analysis signal that User Insights can group for reporting and routing analysis. The signal is not intended to replace or expose the original request. A dedicated Cloudflare Worker processes eligible AI Gateway logs, examining the conversation trajectory — user requests, assistant responses, tool calls, and tool results — to identify the type of work being performed, such as coding, debugging, research, or summarization. It also returns a confidence score and evaluates dimensions including task complexity, intent ambiguity, stakes, and context dependence.

The Worker's category can be joined with the log metadata the dashboard already uses. The same signals support model-fit evaluation by weighing how well candidate models suit a task against their cost. The current implementation deliberately covers a small, understandable set of categories rather than attempting to infer every detail of a user's work.

The pipeline follows the existing AI Gateway log architecture: metadata is stored separately from log bodies, currently using Durable Objects for metadata and R2 for log bodies. User Insights exposes derived categories and aggregate views; it does not turn the dashboard into a raw prompt browser. Retention of underlying log bodies continues to follow configured AI Gateway logging behaviour, so teams should review those settings when deciding what to send through the classifier.

Classification is asynchronous. AI Gateway writes the log to its existing storage path first, and the Worker processes it afterwards, keeping classification off the request path so it adds no latency to the user's response. The tradeoff is that User Insights is not a real-time view: new conversations may not appear immediately, and analysis can trail incoming traffic by roughly one day as logs are processed and aggregated. It is suited to spotting usage patterns over time rather than monitoring live activity.

BLOG-3566 7.png
 User Insights Flow
BLOG-3566 8.png
Task classification by spend

Attributing activity to people and tools

Task categories gain value when they can be sliced by user, team, or application. AI Gateway is identity-aware, so this context requires no separate reporting pipeline. It works for in-house applications as well as developer tools and agent harnesses such as Claude Code, Codex, and OpenCode. Running AI Gateway behind Cloudflare Access lets teams connect authenticated users and sessions to their AI traffic so User Insights can attach activity to the right person and conversation.

Custom applications must include both a stable user_id and a session_id in their requests for User Insights analysis. Exact identity configuration and field names depend on how the application or tool is set up; the key requirement is stable, non-sensitive identifiers so usage can be grouped without placing identity data in the prompt itself.

POST https://gateway.ai.cloudflare.com/v1/$ACCOUNT_ID/$GATEWAY_ID/openai/chat/completions

Content-Type: application/json

Authorization: Bearer $OPENAI_API_KEY

cf-aig-metadata: { user_id: \"user-123\", session_id: \"session-456\", idp_group: \"engineering\", application: \"code-review\" }

The request body holds the model and the messages for the conversation. With Access configured in front of AI Gateway, tools like Claude Code, Codex, and OpenCode inherit the identity context automatically. Cloudflare Access is free for teams of up to 50 users, which makes it an easy starting point.

Getting started

AI usage shifts quickly: models change, workflows evolve, and the right choice for one group may be wrong for another. User Insights identifies where models may be overkill, shows which users and agents drive that usage, reveals the tasks behind it, and lets teams compare the cost of completing the work. It is covered by the AI Gateway User Insights documentation, and AI Gateway itself is available in the Cloudflare dashboard.