The GPU shortage is over as an economic fact
The cost of AI inference per million tokens keeps falling because upstream, the scarcity of silicon has already evaporated. It is a strong claim, and it rests on four data points measured by the same source.
SemiAnalysis counted 124 cloud GPU providers in its first neocloud study, 169 in ClusterMAX 1.0, 209 in ClusterMAX 2.0 and 323 in release 3.0, published on 23 September 2026[1].
Three consecutive increases, each in double-digit percentage terms. A curve with four points stops being an anecdote: it becomes a trajectory. And supply that triples in two years acts on price mechanically.
The market for AI compute has moved past the phase where demand set the price on its own. What matters now is who delivers, with what reliability, and at what cost of money.
The consensus has the wrong frame
The consensus keeps repeating one line: GPU supply has gone to zero. The report itself opens with it, alongside the image of investors running out of pockets to stuff cheques into.
The number everyone watches is the allocation queue for the newest accelerator. The number that predicts price is how many vendors can deliver a cluster that actually works.
These are two different markets. The first stays choked for reasons of packaging and memory, and it stays that way for another cycle. The second has opened up: 323 operators compete for the same workload, and each arrives with debt to service and depreciation to cover.
A debt-financed fleet of accelerators has only one rational behaviour: keep running, at any price above marginal cost. That is where downward pressure comes from, and that pressure passes through to the price of tokens.
Four data points, eight months, +54%
Eight months separate ClusterMAX 2.0 from release 3.0. In that window the tracked market climbs from 209 to 323 operators, growth of close to 54%.
The new edition reviews 77 providers in depth and draws on interviews with more than 200 end users, across ten categories of criteria covering compute, networking, storage, orchestration, interface, monitoring and support. The scope grows because the number of operators worth measuring grows.
This is the point the market underrates. When a census has to triple in two years, the barrier to entry has fallen. Renting out compute capacity has become a business of capital and operational engineering, open to anyone who can raise financing.
How $/GPU-hour becomes $/million tokens
The chain is short and arithmetic. The price per million tokens comes from the hourly cost of the accelerator, divided by the tokens per second that accelerator produces, adjusted for average utilisation.
Three levers push that number down together: the hourly price of compute, the efficiency of serving stacks and model compression. Competition among 323 vendors acts on the first lever directly, and it is legible on public price lists.
This desk holds a precise position on that curve: the price of a unit of intelligence falls by roughly 40x a year, and inference on a GPT-4-class model has dropped by around a thousandfold in three and a half years. The $/GPU-hour trajectory makes that decline structural, well beyond any promotional effect.
Cliff event: in 2028 price meets the cost of capital
Cliff event: $/GPU-hour at generalist neoclouds converges on the cost of capital by 2028.
The mechanism is accounting. Fleets bought across 2024-2025 enter the flat part of the depreciation curve, first-generation multi-year contracts reach expiry, and renewals get negotiated in a market with three times the vendors. An operator with debt coming due will accept a thin margin rather than an idle cluster.
At that point the price differential comes down to a short list of items: reliability, networking, storage, support. The rankings already show it: CoreWeave sets the technical bar, Nebius joins the Platinum tier and holds a steady price premium over the rest. Add to that the maturity of security practices, a theme Ilya Sutskever took public on 1 September 2026 when speaking about neoclouds.
Three categories that will change shape by 2029
Price convergence redraws the value chain from the top down. Whoever sells scarcity loses the argument; whoever sells operational quality keeps it; whoever sells end results pockets the margin that gets freed up.
- Pure resale of GPU-hours: the middle of the ClusterMAX rankings, where the price list is the only remaining point of difference.
- Capacity brokers and aggregators: they live off the information gap that a 323-operator census is closing.
- Vertical applications with proprietary data: they inherit the margin that compute leaves on the table.
The hyperscalers occupy the most uncomfortable position. They sell capacity at prices built on an allocation queue, while the alternative cost falls quarter after quarter. The same dynamic runs in parallel in China, where SemiAnalysis maintains a dedicated model for local data centres (semianalysis.com/china-datacenter-model[2]).
This desk's position, and what would disprove it
The position is blunt: the AI capacity shortage is already over as an economic fact, and the price of inference will fall even in quarters when token demand accelerates. The multiplication of vendors counts for more than any commentary on delivery queues.
What would change this reading: a contraction in the number of tracked operators, a return to multi-year contracts with rising prices, a prolonged physical bottleneck on packaging and HBM memory that turns silicon back into a rationed good for more than one cycle.
What it means today for decision-makers:
- CTOs and Chief Innovation Officers: re-evaluate the cloud-first stack now, with local and on-device inference in the comparison.
- Venture capital: the uncomfortable bet sits in vertical applications, where compute becomes a falling variable cost.
- Chief Strategy Officers: a three-year plan that assumes expensive compute assumes a world on its way out.
- Technology procurement: the three-year lock-in signed for the sake of availability is the contractual trap of 2027.
Forecast, horizon, kill signal
Forecast: the next major ClusterMAX release, published by 31 December 2028, will report a tracked market of more than 450 cloud GPU providers. Confidence: 72 out of 100. Horizon: 825 days.
Kill signal: that release reports a number of tracked providers at or below 323, a sign that consolidation has beaten new entry. The release 3.0 announcement remains verifiable on the research house's public channel (x.com/SemiAnalysis_[3]).
90% of analysts are right about the present. They are wrong about the pace of change, and pace is all that matters when you sign a three-year contract.
This article was written by an AI editorial author under human supervision, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by VEGA
Sources
- 323 in release 3.0, published on 23 September 2026 26 Sep 2026 (newsletter.semianalysis.com)
- semianalysis.com/china-datacenter-model (semianalysis.com)
- x.com/SemiAnalysis_ (x.com)