DeepSeek breaks into production volume; Anthropic still owns the spend

Vercel's AI Gateway production index for May 2026 tells a story of two markets diverging. Total token volume through the gateway grew 20% month-over-month, while total spend climbed 43% — meaning customers paid roughly 20% more per token on average than in April. The split is driven by a new low-cost entrant that captured significant volume without denting the revenue share of frontier labs.

DeepSeek jumped from under 1% to 17% of gateway tokens in a single month, powered by the May release of deepseek/deepseek-v4-flash and deepseek/deepseek-v4-pro. The models launched at prices roughly 20–50x lower than comparable Anthropic offerings and 8–12x below value-tier competitors like Qwen 3.6 Plus and Kimi K2.6. V4 Flash entered at $0.14 input / $0.28 output per million tokens.

In May 2026, DeepSeek held 17% of monthly tokens, putting it third on the gateway by token volume.

The spend data tells a different story. DeepSeek's share of cost landed near 1% even at 17% token share. Teams testing V4 against existing evals found output good enough to ship — price alone wouldn't have driven that volume shift in a month. It marks the first time a model at that price point cleared the quality bar for production work at scale.

DeepSeek was prominent in the previous token volume chart, but is nearly invisible in this spend chart.

Frontier spend grows as high-stakes workloads stay Anthropic-heavy

While the low-cost end grew fastest in volume, the expensive end grew faster in dollars. Anthropic's token share rose from 26% to 32%, and its spend share from 61% to 65%. OpenAI held near 13% token share but ticked up to 13% of spend on a much larger total — customers paid more per OpenAI token in May.

The coding agent segment shows the split most clearly. DeepSeek drove 49% of that segment's tokens but only 4% of its cost. Anthropic drove 28% of tokens and 70% of the cost. Work that demands frontier models grew faster than work that doesn't, which pushed the average token price up despite DeepSeek pulling it down.

In May 2026, DeepSeek took almost half of the coding agent use case, with xAI and MiniMax dropping off significantly. Back-office workloads stayed Anthropic-heavy across both months.

Across every high-stakes use case — AI app generation, back office agents, and coding agents — Anthropic held 70–80% of spend in May, up from 61% overall in April to 65%. Prior headlines about blown token budgets, including Uber burning through its annual Claude Code budget in Q1 and Amazon shuttering its KiroRank leaderboard, didn't stop production spend from increasing.

Anthropic continued to own high-stakes use cases in May 2026, even with DeepSeek V4's significant gain in token volume.

Teams route around price increases

The data shows deliberate routing strategies rather than blanket budget cuts. Teams sent high-volume, lower-risk work to cheap models and reserved frontier models for quality-critical tasks. One clear example: Google's Gemini 3.5 Flash launched in May at a higher price than Gemini 3.0 Flash, and migration stalled. By month-end, 3.5 held only 7% of the Flash family's tokens while 3.0 retained 90%.

When Gemini 3.5 Flash launched in May at a higher price than Gemini 3, migration didn’t happen at scale.

That contrasts sharply with the rapid adoption of Gemini 3.1 Pro in February and March, which jumped to 30% immediately and dominated its family within a month. Teams that were satisfied with 3.0 Flash weren't willing to pay more for 3.5 yet — a sign that upgrades are now measured against ROI rather than taken by default.

When Gemini 3.1 Pro launched in February, it gained 30% adoption immediately, and by the next month was the dominant model in the family.

High-volume work gets cheaper, high-stakes work gets costlier

Two optimization strategies emerged from the May data. First, teams adopted DeepSeek's V4 family for lower-risk, high-volume tasks. Second, they delayed model family upgrades when the price increase didn't justify the benefit. Routing gives teams the ability to adjust their model mix and budget in real time.

Volume-heavy workloads run cheaper per token. Personal assistants and coding agents cost the least, while back-office and recruiting work — where a wrong answer is expensive — cost far more. This reflected in the month's overall shift: B2B applications ran roughly 60% more expensive per token than B2C, since B2B runs fewer but more consequential calls.

B2C drives token volume while B2B drives spend.

Agentic workloads continue to punch above their weight. Just under a quarter of requests end in a tool call, but those requests carry well over half of all tokens, running about 2.5x denser per request. Both metrics were roughly flat month-over-month.

Agentic traffic remains far more token-heavy than its request share suggests, running about 2.5x denser per request on average.

Scale also drives model diversity. Apps serving under a modest request volume typically run a single model, but at the 1M+ request tier, the majority route across 11 or more distinct models.

Model diversity rises with scale. At 1M+ requests, teams route across 11 distinct models or more.

About the data

This analysis is based on anonymized, aggregate routing data from the Vercel AI Gateway through May 2026. Spend uses market-rate published list prices to normalize across teams bringing their own API keys. Volume counts tokens routed through the gateway. B2C, B2B, and use-case classifications are aggregate; no individual team or workload is identified. Prior months may be revised as the most recent data completes.