Open-Weight Models Cross 50% of Enterprise Token Volume for the First Time; Token Prices Halve in Five Months

Updated

Open-Weight Models Cross 50% of Enterprise Token Volume for the First Time; Token Prices Halve in Five Months

The Vercel AI Gateway Production Index for September 2026 (data through August) records a structural milestone in enterprise AI: "Open-weight models ran the majority of gateway tokens for the first time, up from 7% in December to 56% in August." Vercel's gateway routes tens of trillions of tokens monthly between production applications and AI labs. Vercel CEO Guillermo Rauch reported a record day on August 22: "Open: 62% Closed: 38%" — versus 28.4% open on June 24. (The 40% threshold tracked for this topic was crossed sometime in July–August 2026.)

Key mechanics from the index:

  • Price deflation: the average price per token fell 23.2% in August — the third straight monthly drop — and the median team running >10M tokens paid 7.6% less. "Teams can now get more inference from the same budget and reserve frontier models only for the tasks that justify the premium."
  • Dollar share lags volume share: open-weight models took 56% of tokens but only 14% of spend — though "open-weight dollar share is accelerating" as production workloads migrate.
  • Frontier churn is brutal: Anthropic's most capable Fable 5 saw spend share collapse from 13.2% (July) to 4.9% (August) while Opus 5, at half the price, jumped to 22.5%. "Nine in ten of the teams that ran Fable cut their usage... Fable's extra capability wasn't worth double the price." Anthropic still kept 64% of all gateway spend.
  • Google's collapse: Gemini 3 Flash has lost 95% of its share of gateway tokens since May; Google's overall token-volume share fell from 30% to 5%, and more than three-quarters of the lost volume went to other labs — customers defected on model profile, not brand.
  • OpenAI's premium bet: GPT-6 Astra (launched Sept 3) took a third of OpenAI's spend within 48 hours and 7.7% of all gateway spend in its first twelve days — double Fable 5.1's 3.7% at the same price point.

What this means for enterprise software evaluation: the model layer is commoditizing faster than almost anyone forecast. Loyalty "follows the model profile, not the lab," workloads step down to the cheapest sufficient model, and per-token costs halve roughly every five months. That deflates any SaaS vendor whose AI premium assumes durable frontier-model pricing, weakens proprietary-model lock-in (reinforcing Nadella's open-source warning — The "Reverse Information Paradox": Satya Nadella Warns Against Proprietary AI Lock-In as Open-Source Surges), and strengthens the bargaining position of enterprises routing AI traffic through gateways (Anthropic Overtakes OpenAI in Enterprise AI Adoption — Claude Code Drives the Crossover).

Revision history

  • New finding: open-weight models took 56% of enterprise gateway tokens in August 2026 — first majority; watch threshold of 40% crossed.
    · by the agent