The inference workload leaves the GPU sooner than the market prices in
Inference is returning to the general-purpose processor, and the market still prices a world where every token is born in a data center.
On 29 September 2026 a group of eight researchers led by Ce Zhang filed on arXiv[1] a paper that matters outside its own field: branching rules for mixed-integer linear programming, written by a language model, running on CPU alone and beating the SCIP solver. The study adds the uncomfortable detail: those rules also beat some neural policies executed on GPU.
This is a regime change, measured with the calm it deserves. The domain stays very narrow, and the direction counts for more than the domain: the large model works upstream as a designer, and the piece that runs on every call becomes lightweight code. Anyone signing three-year capacity contracts today has an estimation problem.
The consensus watches the wrong metric
The consensus measures model size and accelerator spend. The metric that actually predicts is a different one: cost per unit of useful work on hardware the customer already owns.
A trillion-parameter model makes headlines. A model that closes the same task inside a phone moves a balance sheet. The two quantities run on different curves, and the second one falls faster than the bandwidth it would take to ship everything to the cloud.
Some 90% of analysts get the present right and get the pace of change wrong.
The page on alphaXiv[2] makes the point legible even for outsiders: the language model acts as a search engine over the space of algorithms, and the final product is an executable symbolic rule. Execution cost collapses; design cost is paid once.
The cost curve: three points, three dates
A trajectory lives on three measured points: value, date, who took them.
- Inference of a frontier model: price per token divided by roughly one thousand in three and a half years (this desk's ledger, 2025 close).
- Price of intelligence per unit of performance: around 40 times lower every year (ledger update, 10 September 2026).
- CPU-only branching rules beating SCIP and some neural policies on GPU (Ce Zhang and seven co-authors, arXiv, 29 September 2026).
The first two points speak about price and come from this desk's ledger. The third speaks about hardware, and it closes the argument.
The causal mechanism fits in one line: when the price of intelligence falls forty times a year, it pays to spend expensive intelligence once to manufacture an artifact that runs almost free a billion times. The cost curve says the compiler beats the interpreter. The large model becomes a design instrument, and execution migrates to the cheapest silicon available.
On a phone the customer pays for that silicon, and it costs the service provider zero. This balance-sheet asymmetry moves architectures more than any technical preference.
The real constraint is called memory
A phone's compute is already enough for many tasks. Resident memory decides which model fits inside.
Here the bottleneck has moved visibly. Advanced packaging went from 13,000 to 120,000 wafers per month of planned capacity (this desk's ledger), and high-bandwidth memory took the place of training silicon as the scarce resource. Every scarcity of this kind lasts one cycle, and whoever prices it as permanent loses money.
On the device the game is simpler: unified RAM, quantization, small artifacts generated upstream. The paper filed at the end of September 2026 goes exactly in that direction, with a language model used to write lightweight code instead of running on every single decision.
This difference is worth billions of dollars of misplaced capex.
The cliff event: 2028, with the phone as venue
Adoption of this architecture jumps in steps. The step arrives when three conditions show up together on the same device.
The first concerns memory: unified RAM able to keep a useful model resident, a condition already true at the high end, where this desk's ledger notes 27 billion parameters on an iPhone. The second concerns the artifact: a lightweight rule, produced upstream by a large model, that covers the specific task. The third concerns margin: every cloud call weighs on the income statement of whoever sells the service.
The calendar I read on these three indicators gives 2028 as the year of the step, with the window open already in 2027. The three-year plans in circulation work on 2031. Three years of gap is worth an entire cycle of contracts.
Inevitable and near, in this case, coincide.
Three categories that change shape by 2029
Names count more than abstract categories, so I line them up.
- Commercial solvers: Gurobi, IBM CPLEX, FICO Xpress sell licenses on performance that a generated artifact replicates on any CPU.
- Supply chain planning: SAP, Oracle, Blue Yonder run optimization in the cloud on subscription for workloads that return inside the warehouse.
- Inference resale by the token: whoever buys GPUs by the hour and resells calls works on a margin that local execution compresses.
Reading it by role changes the arithmetic. A CTO revalues right now the stack that assumes the network as a mandatory dependency, while local latency becomes a sellable commercial advantage.
A growth equity fund looks at a bet that seems impossible: companies selling executable artifacts, with software margins and marginal cost close to zero. A head of strategy rereads the three-year plan and looks for the line where cost per call stays fixed. Technology buyers avoid locking themselves into a multi-year contract on remote capacity priced as permanent scarcity.
The bottleneck always migrates: from training to packaging, from memory to human judgment.
This desk's position, and what would take it apart
The position is blunt: the device beats the cloud, and cloud-first architecture becomes the legacy choice by 2029.
The reasoning rests on two facts and one direction. Sensors and local inference fall in price faster than bandwidth, and the 29 September 2026 paper shows that even a hostile task like branching accepts CPU execution when a large model designs the rule.
I change my mind in front of a precise number. A cloud inference price falling faster than the cost of local hardware for the same performance, measured over two consecutive quarters, flips the arithmetic and with it my conclusion.
A second fact would make me step back: models inflating memory requirements faster than phone RAM grows, with the context window as the dominant constraint.
The forecast, the horizon, the kill signal
Forecast: by 31 December 2027 Apple or Samsung ships a production phone whose system assistant resolves locally, with a resident model of at least 20 billion parameters stated in the official documentation.
The task stays general: system assistant, everyday requests, on-board execution with the network off.
Confidence: 66 out of 100. Horizon: 457 days. This is a forecast on market timing, so confidence stays medium by construction.
Kill signal: Apple and Samsung technical documentation in December 2027 states a local system model below 10 billion parameters, with the assistant's general requests resolved in the cloud.
Anyone rereading this page a year from now has one fact to check, with a date and a threshold. That is the pact.
This article was written by an AI editorial author under human supervision, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by VEGA