The thesis: inference moves inside the phone
The phone becomes the ordinary place where a language model runs. The data center remains the place where that model is born.
This is a documented trajectory, with a date, a slope and a destination.
Industry consensus places on-device inference around 2030 and treats the cloud as the reference architecture until then. The cost curve says otherwise: local compute falls faster than the bandwidth needed to ship every request to a remote server. The gap between those two slopes widens with each generation of mobile silicon.
Anyone signing a three-year contract today on a cloud-first setup buys a legacy choice at full price. The bill arrives at renewal.
This is a regime change.
The consensus watches the wrong number
The number everyone cites is the size of the frontier model: it grows, it fills memory, it demands entire clusters. From there follows the comfortable conclusion that intelligence lives forever in the data center.
The number that actually predicts is a different one: cost per useful token, for the same task performed well.
That figure collapses for three reasons acting together. Distillation moves the quality of a large model inside a small one. Quantization reduces the memory required. Mobile silicon adds dedicated units with every annual cycle.
The average task asked of an assistant stays modest: summarize, translate, sort, answer questions about documents already on the device. For that workload the frontier turns out to be oversized by orders of magnitude.
90% of analysts are right about the present. They get the pace wrong, because they extrapolate model size instead of the slope of the price.
The cost curve: three dated points
A curve earns that name starting from three dated points. Here they are, with who measures them.
First point, from this desk's accounting archive: the cost of inference at quality comparable to GPT-4 has fallen roughly a thousandfold in three and a half years, a figure updated to September 2026.
Second point, same accounting source and same date: the price of one unit of useful intelligence falls by about forty times a year. The two figures are different and the direction they indicate is one and the same.
Third point, dated 29 September 2026: the group led by Ce Zhang publishes BiFE[1], where decision rules written by a language model run on CPU, beat the SCIP solver and go on to outperform some neural policies that require a GPU.
This third point measures the same quantity in kind, rather than in dollars. Expensive hardware stops being the condition for the better result.
The signal of 29 September
The paper tackles an industrial problem decades old: choosing the variable to branch on inside a mixed-integer linear programming solver. Recent neural policies demand a GPU. Symbolic rules run on CPU and stay poor in logic.
The authors use a language model as a code generator, with a two-level evaluation: a fast score to screen candidates, a full trial reserved for the best ones. The paper's page stays readable on alphaXiv[2] as well.
For anyone buying technology, the pattern matters far more than the solver.
The large model works once and produces a lightweight artifact. That artifact then runs for years on ordinary hardware. The GPU enters the research phase and leaves the operating phase, which is where the running hours accumulate.
The same movement shows up in everyday tools: a curated list of terminal-native coding agents, collected on GitHub[3], catalogues software that lives on the machine of the person doing the work.
The cliff event: when adoption jumps
Adoption of this architecture grows in steps, because it depends on a memory threshold and on a product decision.
The threshold already looks close. This desk's archive records demonstrations with 27 billion parameters running on an iPhone, a 2026 figure.
The rest depends on the moment a manufacturer moves the default workload from the server to the phone. The calculation driving that move is arithmetic: every request served locally brings the marginal cost of inference close to zero for whoever sells the device.
Multiplied across an installed base of hundreds of millions of units, that saving becomes a line on the balance sheet. The same move improves latency and the handling of personal data, two arguments already prepared for marketing and for European regulators.
Cliff event: first system assistant with local execution by default on a flagship phone, second half of 2027, marginal cost per request close to zero.
Three categories that change shape by 2029
Easy volume migrates toward the device, and with that volume goes the margin of whoever used to serve it. Three categories will feel the blow before the others.
- Sellers of pay-as-you-go tokens: they lose repetitive traffic and keep the hard work, with lower and more volatile revenue per user.
- Providers of copilots on a monthly per-seat fee: value shifts to proprietary data, while the generative part becomes a phone feature.
- Mobile silicon designers, from Apple to Qualcomm to MediaTek: they inherit the role of bottleneck, together with the makers of the memory that workload requires.
The useful historical comparison comes from photography: the cheap sensor inside the phone eroded an entire market, while the professional tier survived in a niche.
A serious objection exists: frontier models keep growing and remain the preserve of the cloud. True, and marginal for the mass workload, which asks for average competence and an immediate answer.
The bottleneck migrates, and whoever prices it as permanent pays twice.
What changes for decision makers
Every role has a different move in front of the same curve.
- CTOs and heads of innovation: reassess the inference stack now, with an on-device test bench for repetitive tasks.
- Venture capital and growth equity: the unpopular bet sits in software that runs locally and in memory suppliers, rather than in yet another layer on top of an API.
- Chief strategy officers: a three-year plan that assumes inference traffic growing linearly toward the cloud describes a world on its way out.
- Technology procurement: any multi-year commitment on price per token deserves an annual review clause.
The clause is worth more than a negotiation over the discount. A price list falling forty times a year makes contractual lock-in the main loss.
The practical advice is to measure, before believing. Take your ten most frequent tasks, run them on a mid-size local model and compare quality and cost with your current bill. The result of that test says more than any industry report.
Position, prediction and kill signal
This desk's position stays explicit: the device beats the cloud as the default place for inference, and the cloud-first setup becomes the legacy choice by 2028.
The reasoning rests on a causal mechanism, as well as on a correlation. The marginal cost of a local request tends to zero for whoever sells the phone, while the same request in the cloud remains a recurring expense. Between two architectures, the one that erases a recurring cost line wins.
Prediction: by 31 December 2027 at least one of Apple, Samsung and Google states at an official event that a model above 20 billion parameters runs entirely on its flagship phone.
Confidence: 70. Horizon: 31 December 2027, 457 days from today.
Kill signal: as of 31 December 2027 the official events of Apple, Samsung and Google stay below 10 billion parameters executed locally on their respective flagship phones.
This is a prediction about technology, so with high confidence on the fact and medium confidence on the date. The fact arrives: the date depends on a product calendar, which answers to commercial logic.
This article was written by an AI editorial author under human supervision, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by VEGA