2026-10-11 17:15 UTC

Bloomberg reports, citing people familiar with the matter, that Harvey's gross margins swung from about +50% in January to โˆ’50% by June 2026 under frontier-model token costs and turned positive again after it shipped a Kimi K3-based custom model โ€” confirmation would make negative application-layer margins a demonstrated driver of professional-agent vendors' shift to open-weight models.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: highinference-economics ai-application-margins open-weight-modelsHarveyGabe PereyraMoonshot AI
Surfaced 2026-09-28T05:38:07Z โ€” Bloomberg's own exclusive reporting: "After a March update to its AI agents, Harvey's customer usage spiked, but its gross margins dropped s โ€” The FT enterprise-pivot story has become a genuinely hot Reddit thread (319 pts/105 comments, ~97th peer percentile, rate ~53 pts/h after a 29x velocity-spike firing, steady momentum, three-platform spread) โ€” a loud amplification wave, not new substance: still no second source on Harvey's margin swing, no post-Tenet number, and comments add only anecdote (one anonymous F50 engineer evaluating on-prem Kimi/Qwen hosting). Heat reprices lowโ†’high because the measured speed and magnitude-valve spread contradict the stale 3 pts/h reading that priced the last 'low' โ€” Scott's flagship unit-economics episode is peaking on platforms right now โ€” while belief stays corroborated and the thread adds nothing material.

What is this?

Harvey, a legal-AI vendor valued around $15.5B, saw its gross margin swing from roughly +50% in January to about โˆ’50% by June 2026 after a March agent update drove ~20x growth in customer token usage that it paid for under OpenAI/Anthropic usage-based enterprise pricing while charging flat per-seat fees. It turned positive again after shipping Harvey Tenet in August, an in-house model post-trained on Moonshot AI's open-weight Kimi K3 (with Fireworks), which Bloomberg reports performs near Anthropic's strongest models at a fraction of the cost. The margin figures come from a single unnamed source with Harvey declining to comment on specifics, and no post-Tenet margin number or share of workload still routed to GPT/Claude has been published; Bloomberg names Abridge, Decagon, Ramp and Rogo as making similar open-weight pivots, backed by Sequoia and General Catalyst.

Why it matters to Scott

Converges with dated receipts: Harvey's flat per-seat book against ~20x metered token burn is fixed-price-is-underwriting failing in public (mispriced underwriting on a usage-open exposure), and the Tenet fix โ€” owning weights post-trained on open Kimi K3 โ€” is context arbitrage plus capability symmetry executed at $15.5B scale, extending the open-weight pivot lineage from dev-tool cost cuts (Coinbase/GLM-Kimi) to vertical professional SaaS margin survival; note K3's API stays near frontier pricing, so the playbook is weight ownership, not cheap API. Flagship receipt for the Mature Token Law / unit-economics publishing line and direct advisory material for LeverageAI on pricing agent offers โ€” with the China export-restriction tail risk making 'own the weights' advice non-trivial; treat the economics claim as provisional (single unnamed source, no post-Tenet margin number).
ip:framework.fixed-price-is-underwritingip:concept.context-arbitrageip:concept.capability-symmetryip:concept.model-perishabilityip:source.the-mature-token-law-ebookip:concept.successor-unitwork:project.leverageairadar:concept.inference-economicsradar:concept.kimi-k3radar:concept.open-weight-modelsradar:open-weight-inference-economicsradar:coinbase-glm-kimi-cost-cutradar:kimi-k3-frontier-pricingradar:china-open-weight-export-restrictions
queries asked of Scott's wikis
  • inference cost unit economics application-layer AI margins
  • open-weight post-training vs renting frontier API tradeoff
  • per-seat pricing vs metered token costs agentic SaaS
  • model routing cascade cheap open model frontier escalation
  • Kimi Moonshot DeepSeek Chinese open models local inference
  • agent token consumption growth long-horizon usage spikes

Measured heat

now 0 pts/hpeak 73 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 506h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-20 14:00โญ origin echo-reconstructedBloomberg's own exclusive reporting: "After a March update to its AI agents, Harvey's customer usage spiked, but its gross margins dropped s
Bloomberg (Rebecca Torrence and Natasha Mascarenhas) on other (echo) ยท attributed from hn.story.49841581
โ€”
09-25 08:08first on hacker news ยท published ยท +114.1hLegal AI Harvey's margins went from +50% to -50% in 6 months on frontier costs
joennlae
โ€”
09-28 00:03first on r/LocalLLaMA ยท published ยท +178.1hFT: Corporate America rejects overpriced frontier, embraces open models
chocolateUI
โ€”
09-25 08:08amplified on hacker newshn.story.49841581
joennlae
peak 3 ยท 2 comments ยท 1% of case engagement
09-28 00:03amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wrzpzg
chocolateUI
peak 681 ยท 208 comments ยท 98% of case engagement
09-29 18:46amplified on hacker newshn.story.49898418
martinald
peak 7 ยท 0 comments ยท 1% of case engagement
09-25 08:21our radar first saw it ยท +114.3hdiscovery anchor: hn.story.49841581โ€”
09-28 05:31reached heat=high ยท +183.5h ยท via ledgerโ€”โ€”
pace: p90 vs 1032 stories at the 336h mark (now 506h old) โ€” ahead of claude-cowork-chat-unification (1.0x), behind debian-llm-usage-vote (1.0x)

Evidence (4) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnLegal AI Harvey's margins went from +50% to -50% in 6 months on frontier costs
Retrieved article excerpt

Open article ยท Retrieved 2026-09-25T08:24:33.153446+00:00

[Bloomberg](https://www.bloomberg.com/company "Bloomberg")

# OpenAI, Anthropic Costs Push More Startups to Build Off Cheaper Open Models

The Kimi K3 webpage on a laptop arranged in Hong Kong. ยท Bloomberg

Rebecca Torrence and Natasha Mascarenhas

Tue, September 22, 2026 at 2:05 AM GMT+10 8 min read

(Bloomberg) -- The $15.6 billion legal startup Harvey built its business around training AI models like OpenAI's GPT-4 to do specialized work for lawyers. Recently, however, soaring artificial intelligence costs have nudged it to rethink its dependence on the AI giants.

Most Read from Bloomberg

- [Paramount to Settle Lawsuits, Paving Way for Warner Bros.](https://www.bloomberg.com/news/articles/2026-09-21/paramount-to-settle-lawsuits-paving-way-for-warner-bros-merger?utm_campaign=bn&utm_medium=distro&utm_source=yahooUS)
- [Qatar Energy Chief Says Bessent 'Wrong' About Hormuz Future](https://www.bloomberg.com/news/articles/2026-09-20/qatar-energy-chief-says-bessent-wrong-about-hormuz-s-future?utm_campaign=bn&utm_medium=distro&utm_source=yahooUS)
- [AI Risk Is Everywhere and It's Making Billion Dollar Funds Nervous](https://www.bloomberg.com/news/features/2026-09-20/ai-boom-is-making-diversifying-investments-tough-for-wall-street?utm_campaign=bn&utm_medium=distro&utm_source=yahooUS)
- [US and China Hail Talks as Positive Ahead of Trump-Xi Summit](https://www.bloomberg.com/news/articles/2026-09-21/bessent-hails-very-successful-china-talks-on-ai-threats-trade?utm_campaign=bn&utm_medium=distro&utm_source=yahooUS)
- [Qatari Premier Says Gulf States Have to Act Together on Iran](https://www.bloomberg.com/news/articles/2026-09-20/qatari-premier-says-gulf-states-have-to-act-together-on-iran?utm_campaign=bn&utm_medium=distro&utm_source=yahooUS)

After a March update to its AI agents, Harvey's customer usage spiked, but its gross margins dropped sharply โ€” from about 50% at the beginning of the year to -50% by June, according to a person familiar with the matter.Now, the startup has joined a growing group of software companies that are embracing open-weight AI, including increasingly capable alternatives from China. These offerings are typically cheaper than proprietary technology from US AI developers and let firms like Harvey create custom models with their own data.

Investors like Sequoia Capital and General Catalyst are backing the trend, which is giving firms a way to save on one of their biggest costs, while granting them more control over their technology โ€” instead of outsourcing it to OpenAI and Anthropic. This push risks cutting into the AI giants' revenue as both gear up for highly anticipated initial public offerings in the near future. And as debate swirls over the need to pace cutting-edge AI, for some startups, it's also becoming a hedge against a future in which the most advanced models could be slowed or restricted.

Harvey released a model of its own in August, powered by China-based Moonshot AI's Kimi K3, which can perform close to Anthropic's best offerings at a fraction of the cost. That launch, plus other tweaks to its AI usage, have made Harvey's gross margins positive again, according to people familiar with the efforts. The company declined to comment on specific financials for this story.

Similarly, healthtech startup Abridge recently announced it was building a custom foundation model for clinical settings, trained on Nvidia's open models. AI customer support startup Decagon said it now flows 80% of queries through its own models. In fintech, startups including Ramp and Rogo are exploring training their own models for the first time. Coding companies including Cursor, now part of SpaceX, and $48 billion Cognition were some of the first AI applications to release bespoke models.

Story Continues

While model-building was previously too expensive for Ramp, it's considering those efforts now after raising $750 million in June because open-weight systems have improved so dramatically, the startup's co-CEO Karim Atiyeh said. "It made absolutely no sense a year ago. It's starting to make a lot more sense now."

As compute costs climb, "it's completely reckless not to be thinking about that and optimizing for it," said Atiyeh, whose firm is currently in talks to bring in more capital.

Not all investors are bought in, however. Matt Kraning, a partner at Anthropic backer Menlo Ventures, argued that custom model-building doesn't make sense for every company, since it comes with specialized talent needs and higher upfront costs. For some, he said, it's merely a marketing tactic.

"Having your own model or not is such the wrong question," he said. "In most cases, it tends to be a lot of cosplay."

The open-weight debate

Until late last year, AI applications' performance mostly depended on the models underneath them, said Harvey president and co-founder Gabe Pereyra. Training those base models was an extremely expensive effort usually left to the biggest AI labs. Any efforts to train custom models weren't so meaningful in comparison, he said, making it important for companies like Harvey to pay up for the best base options.But the rising costs of the underlying models were putting pressure on some of those companies. So far this year, Harvey has seen a twenty-fold increase in AI token usage, according to the company, a spike its founders knew wouldn't be sustainable if they kept relying on the higher-cost models from OpenAI and Anthropic.

While Anthropic and OpenAI have more recently released lower-cost models in response to customer concerns about AI costs, both now charge enterprises for their model usage on top of base subscription fees, a shift that's punished "tokenmaxxing" pushes. Uber, for example, burned through its full-year AI budget by April after encouraging its engineers to maximize their use of Anthropic's Claude Code.

"If you don't optimize your costs by fine-tuning your own model, you are by definition not efficient," said Dr. Lan Xuezhao, the founder and managing partner of San Francisco-based venture firm Basis Set.

She added: "I don't think the company will be fundable" if it doesn't look to build models for its own use.

Anthropic and OpenAI have also moved more assertively into startup territory. This year, both have been hiring aggressively, releasing plug-ins and launching pilot programs across industries including legal, finance and healthcare.

At a recent investor forum at Anthropic's offices, attendees discussed the rise of startups building their own models. Anthropic mentioned companies including Rogo and healthcare startup OpenEvidence on a slide about how to "ride the curve" of value, according to a source familiar with the matter. It also referenced Harvey's internal work with AI, with a reminder that the startup still needs Claude's Opus, one of Anthropic's most advanced models, for its hardest tasks.Still, some startups worry about what could happen if those model makers cut off their access. Only a week after SpaceX finalized its acquisition of Cursor, OpenAI announced it was suspending the coding startup's access to its models, citing past instances of Elon Musk's companies violating OpenAI's terms of service.

Open-weight offers another option. Closed models from OpenAI and Anthropic keep their internal parameters under lock and key, but open-weight ones publicize those parameters for external developers to download and modify in a process known as post-training.

Some, like Kimi K3 developer Moonshot, even offer their own engineers to help customers fine-tune their off-the-shelf models into customized internal ones, according to two people familiar with the discussions.

The challenges ahead

A shift to open-weight models has its sticking points.

Concerns of data privacy and cybersecurity risks have prompted US lawmakers to weigh various restrictions to open-weight models. Both US security agencies and US-based labs have accused Moonshot and fellow Chinese AI developer DeepSeek of distilling the labs' AI, which is the process of training a new, lower-cost model by continuously querying an older "teacher" model.

Rogo's training efforts, which haven't previously been reported, have necessitated discussions with some of its large customers with hesitations about Chinese models, according to the startup's head of product Strib Walker.

AI data provider Mercor's Chief Executive Officer Brendan Foody said his customers' questions about Chinese models usually come from worries over data usage, rather than potential regulatory crackdowns. Some developers argue, however, that the training process for custom models can be so extensive that the model's parameters ultimately don't look anything like when they started. "They really go from being an open-source model to a Rogo model," Rogo's Walker said.

Also, post-training isn't exactly an easy lift. Menlo's Kraning pointed to the competitive market for specialized AI talent, with engineers that can command salaries into the millions of dollars and could be easily poached by top companies like OpenAI and Anthropic.

Plus, he said, companies fine-tuning their own models need large swaths of proprietary data, or else they must be able and willing to pay for that data. Harvey, which can't access its customers' sensitive legal data to train its models, buys that expertise from Mercor instead.

Not all companies that attempt building their own models have found it worth the squeeze. Salespeak, a company that has mostly avoided venture capital, previously announced plans to build its own large-language models. It ultimately ditched the effort after a few months of trial and error. It did not see "major advantages" over using market-ready large language models, such as those offered by Anthropic and OpenAI, according to co-founder and CEO Omer Gotlieb.

For example, Elorian CEO Andrew Dai, who is building visual AI models, said the economics of running open-source models don't always make sense. Downloading open-weight models and managing the computer infrastructure to host the models can be expensive. For some nascent companies with lower traffic, such as his, it may be cheaper to just pay for a closed model and spend money based on usage, he said.

Startups building their own models generally aren't under the impression that they can stop using Anthropic or OpenAI's tech entirely, even when these labs aspire to compete directly with those startups.

Logan Bartlett, a managing director at Redpoint Ventures which is invested in Anthropic, Abridge and Ramp, only expects the push to reduce reliance on Anthropic's models to grow. Still, he says, startups will continue to use the best intelligence available, even when it costs them. "They're not going to cut off their nose to spite their face," he said.

Most Read from Bloomberg Businessweek

- [Hating on Polyester Is Back in Fashion](https://www.bloomberg.com/news/articles/2026-09-18/why-shoppers-are-turning-away-from-polyester-clothing?utm_campaign=bw&utm_medium=distro&utm_source=yahooUS)
- ['Death Sentences': Crafting America's Favorite Countertops Is Killing Workers](https://www.bloomberg.com/news/features/2026-09-14/silicosis-lawsuits-mount-as-workers-seek-billions-in-damages?utm_campaign=bw&utm_medium=distro&utm_source=yahooUS)
- [Fender Is Making Enemies With a Messy Fight Over Its Iconic Strat](https://www.bloomberg.com/news/features/2026-09-17/fender-takes-on-the-electric-guitar-industry-over-stratocaster-copyright?utm_campaign=bw&utm_medium=distro&utm_source=yahooUS)
- [People Hooked on Vapes Try a New Way to Quit: Cigarettes](https://www.bloomberg.com/news/articles/2026-09-18/to-quit-vaping-some-are-starting-to-smoke?utm_campaign=bw&utm_medium=distro&utm_source=yahooUS)
- [Trump's Quest to Be Consequential Puts the Entire World at Risk](https://www.bloomberg.com/news/articles/2026-09-18/iran-war-tariffs-and-trump-s-second-term-quest-to-be-consequential?utm_campaign=bw&utm_medium=distro&utm_source=yahooUS)

ยฉ2026 Bloomberg L.P.

[Terms]
joennlae32
๐ŸŸง echo.other โญBloomberg's own exclusive reporting: "After a March update to its AI agents, Harvey's customer usage spiked, but its gross margins dropped sBloomberg (Rebecca Torrence and Natasha Mascarenhas)โ€”โ€”
๐ŸŸ  redditFT: Corporate America rejects overpriced frontier, embraces open models
LocalLLaMA
chocolateUI681207
๐ŸŸง hnThe AI margin collapse is gathering pacemartinald70

Interpretation history

Decision trace