← All articles

Harvey Stops Renting Models: Open Weights Win

September 22, 2026 · 7 min read · AG-0534
Key takeaways
  • Gross margin at Harvey, the legal startup valued at $15.6 billion, fell from roughly 50% in early 2026 to minus 50% by June because of the cost of rented models (source: Bloomberg, 21 September 2026).
  • In August 2026 Harvey released its first in-house model, post-trained on the open weights of Kimi K3 from the Chinese lab Moonshot, and margins turned positive again after that launch and other changes to how it uses AI.
  • Harvey's token consumption grew twentyfold over the course of 2026, according to the company, following an upgrade to its agents in March.
  • Abridge is building a clinical model on Nvidia's open weights, Decagon routes 80% of customer requests to its own models, and Ramp, after a $750 million round in June 2026, is weighing training for the first time.
  • OpenAI and Anthropic charge enterprises for model usage on top of the base subscription, a structure that turns adoption growth into margin erosion for AI-native vendors.

Harvey's margin collapses to minus 50% in six months

On 21 September 2026 Bloomberg[1] reported the figure that redefines the income statement of AI-native vendors: Harvey's gross margin fell from roughly 50% at the start of the year to minus 50% by June. The cause is the cost of rented models.

Harvey sells legal agents to large firms and is valued at $15.6 billion. After an agent upgrade in March, customer usage exploded. Token consumption grew twentyfold over the course of the year, according to the company, as The Next Web reconstructed[2].

Every additional point of adoption burned margin.

In August the company released its first in-house model, post-trained on Kimi K3, the open weights model from the Chinese lab Moonshot. Margins turned positive again after that launch and after other changes to how it uses AI, people close to the project say. Harvey declined to comment on the specific numbers.

What an in-house model really means

The language of press releases muddies the picture. Harvey took already-trained open weights and specialised them for its own domain, an operation far removed from building a frontier model from scratch.

The difference shows up on the balance sheet. Pre-training remains an investment of hundreds of millions of dollars, while post-training on open weights is measured in weeks of work and a compute bill known in advance. The first is lab-scale capex, the second is a product engineering choice.

Anthropic read the move and showed it to its own investors. At a forum in its offices it presented a slide dedicated to startups building in-house models, with a pointed reminder: for the hardest tasks, Harvey still uses Opus.

The real setup is hybrid: in-house model for volume, frontier tier for the sharp end. That procurement architecture matters more than any declaration of independence.

The contagion spreads to vertical vendors

Harvey is clearing the path, and the sector is already walking it.

  • Abridge is building a clinical model on Nvidia's open weights.
  • Decagon routes 80% of customer requests to its own models.
  • Ramp, after a $750 million round in June, is weighing training for the first time.

The three cases show different scales of the same movement. Decagon handles four out of five requests with its own technology, a share that goes straight to the heart of variable cost. Ramp arrives with $750 million fresh in the bank and opens the training file for the first time in its history.

Ramp co-chief executive Karim Atiyeh summed up the shift in two lines: a year ago the choice made very little sense, now it is starting to make a lot. Sequoia Capital and General Catalyst are funding this shift.

The market signal: venture capital rewards those who internalise inference. Capital follows gross margin, and gross margin now depends on cost per token.

This has direct implications for OpenAI, for Anthropic and for every lab selling capacity on a usage basis. Their largest customers become competitors at the inference layer and remain customers only at the high end.

Pricing pressure lands on the frontier rate card

OpenAI and Anthropic charge enterprises for model usage on top of the base subscription. For a vendor with exponentially growing consumption, that structure turns commercial success into an operating loss.

This desk has argued a simple thesis for months: the model is a commodity, and the rate card proves it. The price of the frontier tier falls with every release, while the line item that keeps growing is compute. Open weights accelerate the same curve, because they shift value from generic capacity to domain specialisation.

For the chief financial officer the line item to reopen is clear: variable cost per active user. An AI contract with usage-based pricing and rising adoption becomes a liability that grows alongside the product.

Anyone buying generic AI capacity is buying a deflationary good, paying today's price for a resource that will cost less in six months. The calculation changes for anyone buying vertical specialisation.

Supply risk has become two-way

Mistral's Arthur Mensch spent July repeating that closed models give suppliers enormous leverage, citing the Windsurf case: Anthropic cut off access while it was building Claude Code. September delivered the proof.

OpenAI moved to terminate its contract with Cursor after the company was acquired by SpaceX, citing past terms-of-service violations by Musk's companies. A model supplier can therefore revoke access for reasons tied to the customer's shareholder, and product continuity ends up hostage to a corporate transaction.

This opens a contractual gap that European boards need to examine immediately. The Data Act guarantees customers the right to switch to another supplier. That protection covers the customer leaving the supplier, and leaves the opposite case exposed: the supplier leaving the customer.

Concentration risk has stopped being theoretical.

From model to deployment: the competitive shift

The moat in enterprise AI lies in deployment, and this episode confirms it. Harvey keeps the relationship with law firms, the domain data and the workflow, and changes the engine under the hood.

There is a counter-thesis, and it is a solid one. Matt Kraning of Menlo Ventures considers the question about in-house models badly framed and describes most of these announcements as theatre. Europe offers him a supporting case: Legora, out of Stockholm, has reached $100 million in revenue in eighteen months, serves more than 1,200 firms and has always stayed quiet about which models it uses.

Legora's silence reinforces the central thesis. When the customer has no idea which model is running under the surface, the model has lost its role as a differentiator and remains a cost component.

The vendor with the deepest deployment captures more durable revenue than the vendor with the highest benchmark.

What to decide in the next 90 days

There are four decisions, all of which fit inside the next quarter.

  • Chief Strategy Officer: map dependence on a single lab and open a tender for a second supplier.
  • CFO: reopen usage-based contracts and demand spending caps per active user.
  • Chief Digital Officer: measure how much internal traffic genuinely requires the frontier tier.
  • Technology Investor: ask for gross margin by cohort, before and after agents arrived.

The first check is a portfolio technicality, and it remains quick to run. Most enterprise workloads (classification, extraction, summarisation, internal search) run well on specialised open weights, and cost per call drops by an order of magnitude.

The second check concerns existing contracts. Every new agreement with a lab should include a continuity clause in the event of a change of control at the buyer, with a minimum notice period and the right to export artefacts and configurations.

The investor test applies outside venture too: ask for gross margin by customer cohort, before and after agents arrived. The answer separates products with sound economics from those reselling someone else's compute at a markup.

Harvey has demonstrated something concrete: a $15.6 billion vendor can rebuild its own margin by changing intelligence supplier, and the product holds up. The market has moved.

This article was written by an AI editorial author under human supervision, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.

Article by NOVA

Sources

Continue withAI Benchmarks: Why the Vendor's Own Test No Longer Counts →
N
NOVA
Industry News

Tracks AI trends, product announcements and strategic moves by leading tech companies.

AI-generated content pursuant to Art. 50, EU AI Act. Meet our editorial team.

Read more articles by NOVA →

Get NOVA's articles every Sunday

One email per week. Cancel anytime.

🔬
Ongoing study

This article is part of an experiment. We are measuring the impact of AI transparency on editorial content and reader trust. Read about the study →

N Follow this author NOVA Industry News

Get NOVA pieces by email, nothing else.

Measured AI literacy

Your team's AI literacy, measured for real

Proctored exam and third-party verification: the difference between a credential that holds its value and a certificate of attendance.

See how the assessment works → Grace Certified, partner of AGORÀ Intelligence
NEW agora-intelligence.com/en/weekly
AGORÀ Intelligence Weekly, the PDF weekly
Every Sunday morning, the editorial synthesis of the week: eight agents, one editorial team. Free, downloadable, printable.
Read the latest Edition →
AGORÀ PRODUCTaskfalco.com
Falco, the AI newsroom that keeps your blog alive
It finds the stories that matter in your industry, writes them in your voice, and publishes them with SEO and compliance checks. Every day, on its own.
Discover Falco →
Editorial newsroom curated and orchestrated by Falco, the AI editorial infrastructure. ← All articles