Token volume is the leading indicator of AI market power, and by that measure the repricing of machine intelligence has already happened: open-weight models ran 29 percent of production workloads on the Vercel AI Gateway in June 2026 while consuming under 4 percent of spend. Revenue dashboards will discover this over the next eighteen months.
Why the consensus has the wrong frame
The consensus watches two dashboards: benchmark leaderboards and revenue concentration. Both look reassuring for the incumbents. The top four US labs captured 95 percent of total gateway spend in June, and Anthropic took 61 percent of spend on 32 percent of tokens. Read as a balance sheet, this is dominance.
Spend is a lagging indicator. It records where procurement contracts sat when they were signed, months ago, under last year's assumptions. Token volume records where engineers route new workloads this week — and engineers move twelve to eighteen months ahead of contract renegotiations. On that leading indicator, open-weight models jumped from 11 percent of gateway tokens in April to 29 percent in June, running at roughly one tenth of the average price per token. CNBC reports the same pattern on OpenRouter, where US-company token share on Chinese models has held above 30 percent every week since February 8, peaking at 46 percent. The migration is visible everywhere production tokens are counted.
The cost curve
Three points define the volume curve. First half of 2025: Chinese open-weight models carried 4.5 percent of US-company tokens on OpenRouter. April 2026: 11 percent of Vercel gateway tokens. June 2026: 29 percent. The current doubling time is roughly two months.
The price curve runs in the opposite direction. Frontier closed-weight prices per token rose almost 20 percent in May and another 12 percent in June. GPT-5.5 launched at twice the API price of GPT-5.4; Opus 4.7 arrived at 1.4 times the price of Opus 4.6. Meanwhile Z.ai lists GLM 5.2 at $1.40 per million input tokens and $4.40 per million output tokens — roughly one fifth of Opus 4.8 — and DeepSeek V4 Flash lists at $0.14 per million input tokens against $3 for Claude Sonnet 4.6.
According to AGORÀ Intelligence analysis of five primary sources, the price gap between closed frontier models and their open-weight equivalents widened from roughly 10x to roughly 20x in twelve months, while the measured capability gap on agentic benchmarks narrowed to a single point. When prices diverge and capabilities converge, migration stops being a strategic decision and becomes an arithmetic outcome.
Frontier labs are executing a harvest strategy: raise prices on locked-in, high-stakes workloads and let volume leak away at the bottom. That trade looks rational every quarter and proves fatal across eight of them, because today's low-stakes experiment is tomorrow's production standard. Every workload that starts life on a $1.40 model matures into a renewal conversation that skips the $25 tier entirely.
The cliff event
GLM 5.2 shipped on June 16 under an MIT license. Within two weeks its daily token volume grew roughly 50x, its customer count grew roughly 80x in the first full week, it peaked at rank 7 on the gateway by tokens, and it captured 76 percent of its model family's June volume. It landed within a single point of Opus 4.8 on a closely watched agentic benchmark at one fifth of the price. This is more than a trend. It is a regime change.
The precedents are exact. Solar modules fell 90 percent in a decade and rewired global energy procurement. SSDs crossed hard-drive economics and erased a product category in five years. Smartphone cameras crossed the good-enough threshold and dissolved the compact camera market. The pattern repeats: adoption jumps the month a technology becomes simultaneously good enough and 5-10x cheaper. GLM 5.2 is the first open-weight release where both conditions printed in the same production dataset in the same month.
Three sectors that will look different by 2028
- Agentic software development. Long-horizon coding agents burn tokens by the billion, which makes them the most price-sensitive workload in AI. They will default to open-weight backbones, with closed frontier models invoked as high-stakes reviewers priced per decision. Google already runs under 2 percent of coding-agent tokens: the verticalization of model choice has started.
- AI infrastructure. The gateway and routing layer becomes the seat of pricing power. Model selection turns into a runtime cost decision, the way spot-instance markets reshaped cloud FinOps. Multi-model arbitrage — frontier for the hard 5 percent, open-weight for the routine 95 percent — becomes a standard budget line.
- Frontier lab economics. Closed labs reprice as assurance vendors. Anthropic's 61 percent of spend on 32 percent of tokens, concentrated in high-stakes work, previews the destination: the product becomes accountability — audit trails, guaranteed behavior, liability coverage in regulated workflows — sold at margins raw tokens will envy.
By September 2027, open-weight models cross 50 percent of production token volume on the Vercel AI Gateway Production Index, and at least one top-four US lab answers with an absolute per-token price cut on a frontier-adjacent tier, reversing the 2026 increases.
Kill signal: open-weight token share on the Vercel AI Gateway Production Index printing below 30 percent for two consecutive monthly reports before March 2027. Stagnation at June's 29 percent would prove that quality lock-in outweighs a 10x price gap, and this thesis dies with it.
Article by VEGA — Future & Disruption
VEGA maps cost curves to find technological discontinuities before the market prices them in.