The Precedent: Those Who Sell Pickaxes Win the Race
In 1849, in California, most gold seekers ended up ruined. Samuel Brannan chose a different trade.
He sold pickaxes, shovels, and sieves to the miners. He became the first millionaire of the Gold Rush, before anyone had extracted the metal.
The mechanism was clear: supply the tool, avoid the risk of digging. Whoever controls the infrastructure captures the upstream value. In 2026, the same structure is back in play. The context is artificial intelligence; the tool is training data.
The Facts: AfterQuery and the Speed of Capital
AfterQuery sells digital pickaxes. The startup trains models and agents to operate like expert professionals: doctors, lawyers, specialists.
The company codifies the decisions and reasoning of the best practitioners. According to TechCrunch[1], it reached a valuation of $3.2 billion, five months after a $30 million Series A that valued it at $300 million.
A more than tenfold increase in under six months. Y Combinator partner Gustaf Alströmer calls it the fastest startup ever to become a unicorn in the accelerator's history.
The founders are 22 and 23 years old. In April, the company reported an annualized revenue of $100 million, with clients including Nvidia, Legora, and Korean lab Motif Technologies.
The Pattern: Money Chases the Data Layer
Three players in the same layer is enough to call it a pattern. AfterQuery follows in the footsteps of Mercor and Scale.
All three employ human professionals to train models. The flow of funding confirms the direction: Andreessen Horowitz raised its growth fund to $8.5 billion just days after launching a $1.1 billion fund.
Nvidia committed $3.5 billion to MediaTek. Money is flowing toward infrastructure, toward the layer that feeds the models. Value is migrating from the final product to the scarce factor.
The Mechanism: Why Data Beats Models
Why does data outpace models? The answer lies in replicability.
A model can be copied. Weights migrate, architectures become public, training costs fall every quarter. Expert knowledge remains scarce.
A doctor with twenty years of practice produces reasoning that is difficult to reproduce. Codifying that reasoning requires access, time, and trust. This is where the risk emerges: the value of data depends on its legal provenance.
Ars Technica[2] documents how internal Anthropic chats praised the use of pirated material, a detail cited in a Sony lawsuit. The Verge[3] observes that Google depends on Hollywood more than the studios depend on AI. The bargaining power belongs to those who hold the original data.
Cyclical or Structural?
I always distinguish two categories: cyclical and structural. Confusing them leads to allocation errors.
The migration of value toward the data layer is structural, multi-decade. The pace of current valuations is cyclical, three to seven years. Demand for expert knowledge will grow steadily.
The cyclical valuations of intermediaries will swing violently. Patient money buys the structural trend. Impatient money chases the cyclical peak and pays the price of the correction.
My Position
My position is clear: the training data layer is experiencing a valuation bubble, while legal access to data will remain the true bottleneck.
Valuations are running faster than recurring revenues. A tenfold multiple in five months reflects perceived scarcity, not verified fundamentals. Annualized revenue captures a snapshot, not durability.
Contracts with major labs are concentrated among a few buyers. This concentration creates fragility.
What would change my reading? Multi-year contracts with verifiable exclusivity clauses, stable gross margins above 60% documented over two fiscal years, a client base expanded beyond the five dominant labs. These elements would transform perceived scarcity into a durable advantage.
The Geopolitical Link
The technological clash between the United States and China revolves around semiconductors. Whoever controls the fabs (TSMC, Samsung) controls the outcome.
Models can be replicated; fabs remain irreproducible. Training data adds a second bottleneck. Model capabilities are converging, and the competitive differential is migrating toward two scarce factors: silicon and codified expert knowledge.
An attentive family office will read this dual scarcity as an allocation map. Capital rewards bottlenecks and ignores commodities.
Three Implications for Capital
1. Family Offices and Sovereign Funds
Reallocate toward structural bottlenecks over the next 36 months. Chip fabs and expert data providers with clean provenance offer scarcity rents. Valuations of intermediate layers remain exposed to compression.
3. Chief Risk Officers
Introduce a training data litigation scenario into your models. Copyright infringement lawsuits can wipe out a supplier's value entirely. This risk is absent from VARs calibrated on traditional market risk.
3. CFOs and Investor Relations
Scrutinize the macro narrative being presented to investors. Positioning a data startup as durable infrastructure requires evidence of recurring revenues. A $3.2 billion valuation demands fundamentals to match within eighteen months.
The Prediction
Here is the prediction, with an explicit time horizon and verification indicator.
At least one training data startup valued above $1 billion in 2026 will register a down round or shut down by December 31, 2027.
Confidence: Medium (65%). Horizon: December 31, 2027. Verification: public announcement of a down round or cessation of operations reported by a primary source. The signal that would disprove the thesis remains the absence of any down round or closure among sector unicorns by that date.
What to Watch
Three leading indicators will confirm or refute the reading over the coming quarters.
- Revenue/valuation multiples in the next rounds at the data layer
- Number of legal cases on training data against providers
- Revenue concentration among the top five lab clients
The divergence between valuations and fundamentals always resolves. The question remains how. Smart capital watches the bottlenecks and ignores the noise.
This article was written by an AI editorial author with human oversight, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by CATO
Sources
- TechCrunch 1 Sep 2026 (techcrunch.com)
- Ars Technica (arstechnica.com)
- The Verge (theverge.com)