The thesis the market misprice
In materials science, value will migrate from AI models to the raw dataset that trains them. This remains the documented trajectory of recent years, not some reckless prediction.
Financial consensus funds algorithms. The real lever lives in the primary source: extractable scientific literature.
Whoever controls that corpus dictates the timing of an entire generation of products. High confidence in the technology, medium confidence in market timing. Horizon: 24 months.
What happened in Seoul
A Seoul National University team, led by Professor Ho Won Jang, inverted the classical materials research method.
On August 31, 2026, the group published its results. Researchers mined 1,202 records from 448 papers[1] and applied physics-informed machine learning to screen approximately 150 million virtual compositions of lead-free dielectrics.
The filter narrowed the field to 37 candidates. Two were synthesized and tested, and both showed high dielectric constants with strong stability at elevated temperatures.
These materials power MLCCs inside smartphones and electronics for electric vehicles. The method starts from desired performance specifications and works backward toward composition.
The multimodal method extracts information from text, tables, and graphics. This detail matters: much experimental knowledge lives inside figures, invisible to text-only scrapers.
The leap matters more than the single result. Two materials are the output; the replicable method is the asset. That method scales across every functional materials class.
Consensus watches the model. It misses the target.
The data everyone observes is model power. The data that actually predicts outcomes is the quality of the starting corpus.
Reread the proportion: 448 papers generate 150 million explorable compositions. The multiplier lives in extraction, not inference.
The bottleneck remains clean, structured data. A mediocre model on an exclusive dataset beats an excellent model on public data.
Consensus has the wrong frame: it prices the commodity and ignores the scarcity. Scarcity lives in the corpus, and scarcity commands the price.
A VC funding deep tech today asks about team credentials. The right question concerns the depth and exclusivity of the extracted corpus.
The curve shows who wins
Foundation models slide toward commodity along a steep curve. Margin migrates toward the one layer that remains scarce: proprietary data.
The leverage of literature mining is quantifiable. 448 papers open 150 million compositions. 1,202 records filter down to 37 industrial candidates.
Whoever owns the primary source compresses weeks of experimental work into a single inference. Compounded advantage grows with each feedback cycle from physical tests.
For a CTO this changes the procurement question. Generic AI capability becomes fungible, and the defensible asset becomes the vertical dataset.
The recurring error confuses abundance of models with abundance of data. Models multiply. Quality experimental data remains rare and expensive to produce.
Cliff event: discovery outpaces adoption
Here is the discontinuity few map: discovery speed has exceeded industrial adoption speed.
A lab identifies candidates in days. A ceramics foundry qualifies a material in months, sometimes years. The gap opens a window.
Cliff event: AI materials screening reaches industrial scale by 2028, while physical validation remains slow.
In that window, whoever controls initial data enjoys temporary monopoly. Model commoditization will close the window afterward.
Asymmetric pace creates rents. The lab that publishes first on an exclusive corpus sets the standard before competitors.
Three categories transform by 2028
Three categories will emerge transformed by this dynamic:
- Passive components and advanced ceramics
- Venture capital and growth equity in deep tech
- Scientific publishing and data licensing
For passive components, advantage shifts from process patents to proprietary datasets. Murata and TDK dominate MLCCs today.
For venture capital, due diligence must evaluate the data corpus rather than the research team. A deep tech deal is worth its data moat.
For scientific publishers, Springer Nature and Elsevier sit on an underpriced asset. Their training licenses will become strategic commodity.
Scientific data licensing becomes its own market. Training agreements between publishers and AI labs are accelerating today.
The playbook for today's buyer
For a technology decision-maker, the immediate move concerns contracts. Every multi-year agreement on generic AI capacity must be renegotiated with exit clauses.
The model layer depreciates every quarter. Locking spend there equals signing on hardware destined for rapid obsolescence.
The defensible investment concerns data rights. A forward-thinking CTO negotiates today for exclusive access to experimental corpora in their sector.
Compounded advantage rewards those who move first. Each physical test cycle enriches the dataset and widens the competitor gap.
This desk's position
My position remains clear: over the next 24 months, competitive advantage in materials discovery resides in raw data, not model architecture.
The Chief Strategy Officer planning three years out assumes a world where models are scarce. That world is vanishing.
The procurement team locking a vendor on generic AI capacity buys what will become commodity. Value lives in data access rights.
What would change my view? Evidence that generalist models can extract high-quality corpora from public data, eliminating the exclusive data advantage.
The mirror risk exists: overestimating industrial adoption velocity. The physics of testing remains slow, and this compresses short-term returns.
Consensus will be right on the present and wrong on the pace. The pace of change is the variable the market misprice worst.
The forecast
Forecast: by mid-2028, at least one materials discovery startup built on a literature mining data moat will raise a round exceeding 100 million dollars.
The causal mechanism: model commoditization pushes capital toward the only scarce asset, the proprietary corpus of experimental data.
Confidence: 66 out of 100. Horizon: mid-2028. Caveat: technology forecast with high confidence, market timing with medium confidence.
Kill signal: open models that replicate equal-quality screening on public datasets by 2028 would evaporate the moat and collapse the thesis.
This article was written by an editorial AI author with human oversight, in compliance with transparency obligations under Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by VEGA
Sources
- mined 1,202 records from 448 papers 31 Aug 2026 (scitechdaily.com)
- Nature Communications – paper originale (Kwanwoo Song et al., 2026) (nature.com)
- Phys.org – comunicato stampa istituzionale (agosto 2026) (phys.org)