What the September index shows
AI Gateway routed tens of trillions of tokens between production applications and AI labs over the period covered by the September 2026 Production Index, which reports on data collected through August 2026. The four headline movements: open-weight models crossed into the majority of token volume, average price per token fell sharply for a third month, Fable 5 gave up most of its spend share to Opus 5, and Gemini 3 Flash continued its collapse in volume share.
A special section covers OpenAI's GPT-6 Astra, which launched September 3.
Open-weight models pass 50% of token volume
Open-weight models ran 56% of all tokens on the gateway in August, the first month they took a majority of volume. In December 2025 they processed fewer than one in ten tokens; token share climbed every month from April through August, from 13% to 56%. That total exceeded the combined volume of all closed-weight models.
Spend tells a different story: open-weight models accounted for 14% of August dollars, though that share is accelerating as capability improves and customers move more production workloads over.
The adoption shift contributed to a 23.2% drop in average price per token in August — the third consecutive monthly decline and the steepest since April. Among teams running more than ten million tokens in both months, the median team paid 7.6% less per token, compared with July's 2.9% decline.
The practical effect is more inference per budget, with frontier models reserved for work that justifies the premium.
Fable's spend share drains into Opus 5
Fable is Anthropic's most capable model; Opus is the tier below it at roughly half the price per token. When the US export control on Fable 5 was lifted and access restored on July 1, its gateway spend share surged to 13.2%. Later that same month, Opus 5 came online. In August, Fable 5's share fell to 4.9% while Opus 5's rose to 22.5%. Nine in ten teams that ran Fable cut their usage, and more moved workloads to Opus 5 than to any other model.
Those workloads stayed inside Anthropic rather than moving to another lab. The company has taken at least 61 cents of every dollar spent through the gateway every month since December, reaching 64 cents in August, and its models have held the top two spots by spend every month since December even as the models occupying those spots changed. Google and OpenAI have each reached third place twice.
Switching behavior follows fit, not brand
Whether a lab keeps its customers depends on whether a new model preserves what users valued in its predecessor. Opus 5 gained almost twice what Fable 5 lost, because it handled the same workloads at half the price. Z.ai's GLM-5.3-Flash, five days after launch, was running three times GLM-5.2's daily volume; it overtook GLM-5.2 one day after appearing on the gateway and processed two-thirds of Z.ai's tokens by August 31.
Google's experience went the other way. Its newer Flash models offered no relative advantage in capability or price, and a majority of Gemini 3 Flash's workloads moved to OpenAI, Anthropic and DeepSeek. Google's own full-size Flash successors took under a tenth of the lost volume; more than three-quarters went to other labs.
That volume split roughly in half by price. One half went to cheaper models, led by GPT-5.6 Luna at less than half the cost per token. Most of the other half went to higher-priced models, led by Claude Opus 5 and Sonnet 5, at roughly nine and three times Gemini 3 Flash's price respectively. Google's share of gateway token volume fell from 30% to 5% over the period, with Gemini 3 Flash accounting for 22 of the 25 percentage points lost.
Gemini 3 Flash's decline
Gemini 3 Flash has lost 95% of its share of gateway tokens since May.
Astra's first twelve days versus Fable 5.1
GPT-6 Astra launched on the AI Gateway on September 3 at the same price as Fable 5.1 and two and a half times the price of GPT-5.6 Sol. Two days later it accounted for one in every three dollars spent on OpenAI models through the gateway, and its spend share has held between 28% and 39% since.
Within OpenAI's lineup, Astra and Sol processed 27% of the company's tokens but accounted for 71% of its spending from September 4 through 16. Luna and Nano processed more than twice as many tokens for about one-ninth as much spending; Luna alone processed more than eight times as many tokens as Astra while Astra accounted for more than four times as much spend.
Anthropic launched Fable 5.1 on September 1, two days before Astra. Over each model's first twelve days on the gateway, Astra took 7.7% of all gateway spend against Fable 5.1's 3.7% — more than double — and was used by twice as many teams. Astra passed Fable 5.1's cumulative gateway spend on day four and reached twice Fable 5.1's twelve-day total by day twelve.
Cheap models carry OpenAI's volume; Astra's early lead over Fable 5.1 shows it can also attract teams at the highest price point. Between the two, OpenAI competes with other frontier labs for both scale and premium spend.



