Open-weight models now carry nearly a third of AI Gateway traffic
Vercel's AI Gateway Production Index for July 2026, based on data collected in June, shows token volume across the gateway grew 29% month over month while spend grew 27%. The price per token held flat, a notable shift after rising almost 20% in May.
The flat pricing is the result of two opposing forces. Open-weight models have climbed from 11% of gateway token volume in April to 29% in June, running on under 4% of spend. That alone should have driven the average token cost down. But closed-weight frontier prices rose about 12% per token over the same period, offsetting the effect. The net result is evidence of the routing discipline seen in prior months showing up in aggregate: high-volume, low-risk workloads go to cheap models, while frontier models handle the high-stakes work.
DeepSeek closes in on Google; GLM 5.2 debuts fast
Open-weight models ran 29% of gateway tokens in June at roughly one twenty-fifth of the dollar cost. DeepSeek is the main driver of that growth, reaching 22.6% of token volume, placing it third behind Anthropic and Google. Google's share slipped to 24% as its April surge unwound, leaving DeepSeek less than two points behind. On current trajectories, an open-weight lab will soon be the second-largest source of tokens on the gateway.
The newest entrant shows the pattern accelerating. Z.ai released GLM 5.2, an MIT-licensed model aimed at long-horizon agentic work, at roughly a fifth of Opus 4.8 pricing. From its June 16 API availability through month end, daily token volume grew about 50x. It ranked #11 on AI Gateway by tokens in its final week of June, hitting as high as #7 on single days. Within two weeks it took 76% of its family's June tokens; Gemini 3.1 Pro, previously the fastest in-family migration on record, took until its second month to reach a similar share.
Frontier labs keep the spend
The top four US frontier labs took 95% of AI Gateway spend in June. Anthropic alone took 61% of spend on 32% of tokens, down from 65% in May but in line with April. It captured 72% or more of spend in high-stakes categories like coding agents, back-office agents, and app generation.
OpenAI's position moved in the opposite direction of DeepSeek's. Its token share fell from 12.5% to 10.3%, while its spend share rose from 13.3% to 16.1%, pushing its cost per token up about 50% relative to the market in one month. Customers are sending OpenAI less volume but more expensive work, a divergence visible only in production data.
Different modality, different leader
Image generation is a two-lab race. OpenAI's GPT Image family generated 53% of images and took 52% of image spend. Google's Nano Banana generated nearly 39% of images on 43% of spend. No other family cleared 5% of either metric.
Video is the one category where Chinese labs hold the premium tier. xAI's Grok Imagine generated 42% of videos on just 19% of spend, occupying the same volume-discount position DeepSeek holds in text. ByteDance's Seedance led video spend at 49% while generating only a third of videos. Together, Seedance, Kling, and Alibaba's Wan took roughly two-thirds of video dollars. These shares are expected to fluctuate as the category matures.
The overall picture requires multi-lab routing: Anthropic leads text, OpenAI leads images, xAI leads video volume, and ByteDance leads video spend. No single provider leads every modality.
Export control halts Claude Fable 5 days after release
Anthropic released Claude Fable 5 on June 9 and adoption was immediate: within four days, the model drew 22 requests for every 100 sent to Opus 4.8. On June 12, a US export-control directive took effect covering the model, and Anthropic suspended access to comply. Fable 5 stayed offline until controls lifted on June 30, with access resuming July 1.
Other notes from June's data
- Back-office agents are the most expensive workload per token, running 5% of total tokens but consuming 14% of total spend.
- B2B use cases drove 46% of June tokens and 60% of spend; B2C was the reverse at 43% of tokens and 26% of spend.
- Google's share concentrates in consumer workloads. It ran 57% of personal-assistant tokens and 54% of education tokens, but under 2% of coding agent tokens.
The analysis draws on anonymized, aggregate routing data from the Vercel AI Gateway through June 2026. Spend figures use market-rate published pricing for a normalized view across teams; volume counts tokens routed through the gateway. Image and video figures count successful media generation requests, not tokens. Open-weight shares are measured conservatively by provider — counting only the four labs that serve their own models on the gateway — while the enterprise-adoption figure classifies by model regardless of provider.



