Open-weight models now carry nearly a third of AI Gateway traffic

Vercel's AI Gateway Production Index for July 2026, based on data collected in June, shows token volume across the gateway grew 29% month over month while spend grew 27%. The price per token held flat, a notable shift after rising almost 20% in May.

The flat pricing is the result of two opposing forces. Open-weight models have climbed from 11% of gateway token volume in April to 29% in June, running on under 4% of spend. That alone should have driven the average token cost down. But closed-weight frontier prices rose about 12% per token over the same period, offsetting the effect. The net result is evidence of the routing discipline seen in prior months showing up in aggregate: high-volume, low-risk workloads go to cheap models, while frontier models handle the high-stakes work.

Usage and spend rose in June; the price per token, up almost 20% in May, held flat

DeepSeek closes in on Google; GLM 5.2 debuts fast

Open-weight models ran 29% of gateway tokens in June at roughly one twenty-fifth of the dollar cost. DeepSeek is the main driver of that growth, reaching 22.6% of token volume, placing it third behind Anthropic and Google. Google's share slipped to 24% as its April surge unwound, leaving DeepSeek less than two points behind. On current trajectories, an open-weight lab will soon be the second-largest source of tokens on the gateway.

DeepSeek reached 22.6% of gateway tokens in June, less than two points behind Google

The newest entrant shows the pattern accelerating. Z.ai released GLM 5.2, an MIT-licensed model aimed at long-horizon agentic work, at roughly a fifth of Opus 4.8 pricing. From its June 16 API availability through month end, daily token volume grew about 50x. It ranked #11 on AI Gateway by tokens in its final week of June, hitting as high as #7 on single days. Within two weeks it took 76% of its family's June tokens; Gemini 3.1 Pro, previously the fastest in-family migration on record, took until its second month to reach a similar share.

GLM 5.2's customer count grew far faster than its token usage in the two weeks after launch

Frontier labs keep the spend

The top four US frontier labs took 95% of AI Gateway spend in June. Anthropic alone took 61% of spend on 32% of tokens, down from 65% in May but in line with April. It captured 72% or more of spend in high-stakes categories like coding agents, back-office agents, and app generation.

OpenAI's position moved in the opposite direction of DeepSeek's. Its token share fell from 12.5% to 10.3%, while its spend share rose from 13.3% to 16.1%, pushing its cost per token up about 50% relative to the market in one month. Customers are sending OpenAI less volume but more expensive work, a divergence visible only in production data.

Anthropic took 61% of June spend; DeepSeek, prominent in the volume chart, is negligible here

Different modality, different leader

Image generation is a two-lab race. OpenAI's GPT Image family generated 53% of images and took 52% of image spend. Google's Nano Banana generated nearly 39% of images on 43% of spend. No other family cleared 5% of either metric.

Two families generated 92% of the gateway's June images: OpenAI's GPT Image models at 53% and Google's Nano Banana models at 39%

Video is the one category where Chinese labs hold the premium tier. xAI's Grok Imagine generated 42% of videos on just 19% of spend, occupying the same volume-discount position DeepSeek holds in text. ByteDance's Seedance led video spend at 49% while generating only a third of videos. Together, Seedance, Kling, and Alibaba's Wan took roughly two-thirds of video dollars. These shares are expected to fluctuate as the category matures.

xAI's Grok Imagine generated the most videos on 19% of video spend; ByteDance's Seedance led spend on 34% of the volume

The overall picture requires multi-lab routing: Anthropic leads text, OpenAI leads images, xAI leads video volume, and ByteDance leads video spend. No single provider leads every modality.

Export control halts Claude Fable 5 days after release

Anthropic released Claude Fable 5 on June 9 and adoption was immediate: within four days, the model drew 22 requests for every 100 sent to Opus 4.8. On June 12, a US export-control directive took effect covering the model, and Anthropic suspended access to comply. Fable 5 stayed offline until controls lifted on June 30, with access resuming July 1.

Over its first four days, Fable 5 drew 22% of Opus 4.8's requests

Other notes from June's data

  • Back-office agents are the most expensive workload per token, running 5% of total tokens but consuming 14% of total spend.
  • B2B use cases drove 46% of June tokens and 60% of spend; B2C was the reverse at 43% of tokens and 26% of spend.
  • Google's share concentrates in consumer workloads. It ran 57% of personal-assistant tokens and 54% of education tokens, but under 2% of coding agent tokens.

The analysis draws on anonymized, aggregate routing data from the Vercel AI Gateway through June 2026. Spend figures use market-rate published pricing for a normalized view across teams; volume counts tokens routed through the gateway. Image and video figures count successful media generation requests, not tokens. Open-weight shares are measured conservatively by provider — counting only the four labs that serve their own models on the gateway — while the enterprise-adoption figure classifies by model regardless of provider.