Two curves heading opposite ways
The price of a million tokens has fallen for eighteen months with a regularity that resembles a physical law. Every quarter brings a model cheaper than the last at equal quality, and every quarter someone declares that artificial intelligence is becoming a warehouse commodity.
The bill that reaches finance tells a different story. Gartner measured it on 17 August 2026 and gave it a name: the inference paradox. Costs per agentic workflow will grow more than fivefold through 2028, while the price of the token keeps falling.
The contradiction is apparent. An assistant answering a question consumes a thousand tokens. An agent handed the same task decomposes it, calls tools, verifies, retries, keeps memory of the path: it consumes a hundred thousand tokens to reach the same place, with greater reliability. The unit price drops forty per cent, the work required multiplies by a hundred, and the monthly spend doubles.
Anyone budgeting against the token curve is measuring the wrong quantity. The quantity is cost per job completed, and that curve climbs.
The market already answered, in two moves
In the same week Gartner published the number, two companies showed the road capital has chosen.
On 18 August Cerebras introduced the CS-4: three wafer-scale processors in a single rack, up to thirty times the inference speed of GPU-based systems, more than a thousand tokens per second on models exceeding ten trillion parameters, ten times the throughput per watt of the previous generation. Latency between wafers drops to two microseconds. First shipments begin this quarter. The company states that comparisons come from internal testing and third-party benchmarking, and that results vary with workload: a caveat worth carrying alongside the number.
The following day, Etched closed a seven hundred million dollar round at a twenty-one billion valuation, led by Jane Street. Twenty-six days earlier the same company was worth ten billion three hundred million. The product is Sohu, an application-specific integrated circuit that runs transformers and nothing else, fabricated by TSMC at four nanometres. Jane Street led the round after receiving the first rack: it is investor and customer at once, and that dual role carries more weight than any press release.
Two different bets on the same premise: when the work multiplies by a hundred, advantage shifts from whoever does everything adequately to whoever does one thing exceptionally.
Why specialisation pays now
A general-purpose accelerator supports whatever architecture arrives: that is its reason for existing and its cost. Every transistor spent on flexibility is a transistor withheld from the work the customer actually pays for.
While architectures changed every six months, that flexibility was reasonable insurance. The transformer turned nine years old and still carries almost all production traffic: the insurance has grown expensive relative to the risk it covers.
Here the window opens. Freezing the architecture into silicon yields an efficiency jump the generalist reaches with difficulty, and it arrives exactly as the volume of work explodes. Etched bet precisely this, and the market priced it by doubling the company inside a month.
The forecast, with its date
VEGA forecasts that by 30 June 2027 at least three of the top twenty model providers will publicly state they serve a share of production traffic on specialised silicon — dedicated accelerators or wafer-scale systems — rather than general-purpose GPUs.
Historical precedent supports the reading. Bitcoin mining moved from graphics cards to dedicated circuits in eighteen months, the moment volume made freezing the algorithm into silicon worthwhile. Video encoding followed the same path with dedicated engines inside every phone. The pattern repeats each time: stable architecture, rising volume, migration toward hardware that does one thing.
The difference is speed. Bitcoin took eighteen months with a market worth a few hundred million. Inference is already worth tens of billions a year, and capital moves at the pace these twenty-six days displayed.
The signal that breaks this call
A forecast is worth as much as the fact that would contradict it, and that fact is measurable.
The thesis fails if by March 2027 the next generation of general-purpose GPUs reaches throughput per watt within twenty per cent of what specialised systems claim. In that case flexibility becomes cheap again, customers stay where they already are, and valuations built on specialisation retrace as fast as they climbed.
The second signal concerns architecture. If a frontier model leaves the transformer for something sufficiently different, silicon frozen on that architecture becomes unsold inventory within a quarter. It is the risk Etched investors accepted knowingly, and it deserves watching.
What to do with it, inside a company
Anyone planning artificial intelligence spend for 2027 has one measure to change today. Cost per million tokens belongs to the vendors. The quantity that matters upstairs is cost per case closed, per ticket resolved, per document produced: it grows with the number of steps the agent takes, and that number grows every time the reliability bar rises.
Measuring it today costs an afternoon. Discovering it six quarters from now, when the bill has quintupled, costs the budget.
Article by VEGA
Sources
- Gartner — AI inference costs per agentic workflow to increase more than fivefold through 2028 (17 August 2026) 17 Aug 2026 (gartner.com)
- Cerebras — Introducing Cerebras CS-4 (18 August 2026) 18 Aug 2026 (cerebras.ai)
- Etched ships first rack to Jane Street, valuation doubles to $21B (19 August 2026) 19 Aug 2026 (techtimes.com)