DeepSeek overtakes Google in gateway token volume

Vercel's AI Gateway Production Index for August 2026, based on traffic routed through July, shows DeepSeek has become the second-largest lab by token volume, running more than twice Google's volume. In April, Google carried nearly 40% of gateway token volume and DeepSeek less than 1%; by July, DeepSeek's share reached a quarter while Google fell to 11%. DeepSeek's cheapest model, V4 Flash, ran more tokens alone than all of Google combined.

The shift is concentrated in consumer-facing workloads. Google's share of personal-assistant tokens dropped by more than half in a single month while DeepSeek's more than tripled, with most of Google's lost share going to DeepSeek. V4 Flash was the most-used model on the gateway in July, accounting for nearly a fifth of total tokens and 70% more than the next most-used model.

Open-weight models finally take revenue share

Open-weight labs have steadily gained volume over the past several months while capturing almost no spending. That changed in July. Open-weight share of gateway token volume grew to 36%, and spend share more than doubled to nearly nine cents of every gateway dollar — the highest in the index's history. Prior to July, open-weight spend had stayed under four cents per dollar from April through June.

The four largest frontier labs' combined share of token spend fell to 89%, its first drop below 93% in seven months, with Google accounting for most of the decline. Almost none of Google's lost spend went to DeepSeek, whose share of spending barely moved even as volume grew. Instead, Z.ai and Moonshot's latest models — GLM 5.2 and Kimi K3 — drove the entire open-weight spend increase. These are the first open-weight models to run meaningful shares of workloads historically held by closed-weight labs.

Kimi K3 ramps quickly in agent workloads

Moonshot released Kimi K3 on July 16, built for long-horizon agent work like Z.ai's GLM 5.2 from June. K3 used about twelve times the tokens per request of its predecessor K2.5 from the start. Daily volume tripled between launch week and the final week of July, and by month's end its requests were as heavy as Claude Opus 4.8's. On the last full day of July, K3 ranked eighth on the gateway by token volume.

The growth was new demand rather than cannibalization. K3 processed nearly two-thirds of all Kimi tokens within two weeks and 82% by the final week, while the rest of the Kimi family's volume fell only slightly. Moonshot's share of total gateway spend quadrupled to 2.3%, and open weight's overall spend share more than doubled to 8.6%, with more than 90% of that growth coming from Moonshot and Z.ai. Kimi K3 and GLM 5.2 are capturing significant volume at more than eleven times DeepSeek's rate per token.

Average price per token falls 13.6%

Companies bought more inference in July and ran more of it on cheaper models. Volume grew 59%, spend grew 37%, and the average price paid per token fell 13.6%. Had June's model mix been held constant, the average price would have stayed essentially flat.

The decline came entirely from routing choices. Among teams running more than 10 million tokens in both months, three in four changed at least a tenth of their model mix and three in five changed at least a quarter. The median team's cost per token fell only 2.9%, but only one team in six stayed within five percent of its starting point. A quarter of teams cut cost by more than 30% — typically those that began July paying well above the gateway average — while another quarter paid at least 20% more, typically teams that had been below average.

OpenAI is a telling example: the average cost of an OpenAI token fell to 58% of its June level because 85% of the volume it added went to GPT-5-Nano, the cheapest model in the GPT-5 family. Overall, 81% of July's tokens ran on models that did not exist on the gateway six months ago. Open-weight models crossed a third of all volume at about a seventh of frontier rates. Google fell from 24.0% of gateway token volume to 10.7%, Anthropic slipped two points to 29.8%, and OpenAI's share rose to 12.8%.

Anthropic's premium widened amid falling prices

Anthropic has held more than 60% of gateway spend in every month measured, even as open weight tripled its volume share. In July it collected 65.1% of all spend on 30% of total volume, with its average price per token running 4.4 times the average of every other lab — up from 3.4 in June.

Part of the increase traces to Claude Fable 5, which returned July 1 after a three-week export-control suspension and quickly grew to 13.2% of all gateway spend, second only to Opus 4.8. In coding agents, the gateway's largest use case by tokens, Anthropic collected more than 80% of spend. Nearly all of Fable 5's July users — nine in ten teams — had not used it before the suspension, even though daily volume returned almost exactly to pre-ban levels.

Anthropic has no model at the bottom of the market. Its least expensive model, Haiku 4.5, runs just over two-thirds of the gateway average price per token; GPT-5-Nano runs a sixth and DeepSeek's V4 Flash a sixteenth. Since switching models on the gateway is a one-line change, customers cutting costs on Anthropic must leave Anthropic entirely, and in July it still took a majority of all spending. The rest of the market's tokens averaged less than a quarter of Anthropic's price, yet Anthropic took two of every three dollars spent.

Media leaderboards change hands

In images, Google's Nano Banana took the lead from OpenAI's GPT Image, generating 45% of the gateway's images to GPT Image's 42%, with spend split almost evenly. Nearly all of Nano Banana's gain came from a single model, Gemini 3.1 Flash Lite Image. Google took the image lead in the same month it fell to fourth by token volume.

In video, ByteDance's Seedance led in both videos generated and dollars spent. xAI's Grok Imagine, June's volume leader, fell to second. Chinese labs collected about seven of every ten dollars spent on video generation.

Additional July findings

  • Within coding — the gateway's largest workload by tokens — DeepSeek ran nearly a third of volume while Anthropic collected more than four of every five dollars.
  • Back-office agents remain the most expensive work per token, with a spending share roughly two and a half times their volume share.

About the data

This analysis draws on anonymized, aggregate routing data from the Vercel AI Gateway through July 2026. Volume counts all tokens routed through the gateway; spend values every request at the lab's published list price; price per token is total spend divided by total volume. Open-weight shares count the four labs serving their own models at scale: DeepSeek, MiniMax, Moonshot, and Z.ai. Image and video figures count media generated, not requests or tokens. Prior months may be revised as methodology is updated.