By June, Harvey's seat price no longer covered Harvey's model bill, and the margin recovered after the company changed the model, not the price. After a March agent update, customer usage spiked and gross margin fell from about 50% at the start of the year to roughly -50% by June, then turned positive again after Harvey launched its own model in August, post-trained on Moonshot AI's open-weight Kimi K3, Bloomberg reported via The Next Web. If you renew any per-seat agentic AI product — legal, clinical documentation, customer support, finance research — your vendor runs on the same arithmetic and has three ways out: meter your usage, cap it, or change the model doing the work. Most enterprise contracts negotiate the first and ignore the other two.
Abridge, Decagon, Ramp, Rogo and Canva are making or weighing versions of the same move. In August we covered Harvey's model as a provenance problem — whose weights sit under your privileged documents. This is the economics half: who pays when an agent uses twenty times the tokens, and what the vendor changes when the answer is "not the customer."
Why Did Harvey's Margin Go Negative?
Harvey's margin went negative because it sells a fixed annual seat and buys a metered input, and agents broke the ratio between the two. Metronome's pricing index, updated in January, describes Harvey as per-seat billing with unlimited usage within purchased seats. The partner who runs one query a week and the associate who leaves a diligence agent grinding through a data room overnight pay the same. Harvey's model providers bill by the token.
Then the product changed shape. A March agent update sent customer usage up, and Harvey has seen a twentyfold rise in token usage this year, according to the company. Over roughly the same period, annual recurring revenue about doubled, from $190 million at the start of 2026 to more than $400 million. Revenue doubled. Tokens rose twentyfold. No flat price survives that ratio.
The swing is worth doing as arithmetic once. At a 50% gross margin, every dollar of revenue carries 50 cents of cost to serve. At -50%, it carries $1.50. The cost of serving each revenue dollar tripled in about six months — and Harvey started from a strong position. Bessemer's 2025 benchmark put the fastest-growing AI "Supernovas" at roughly 25% gross margin, often negative, and the steadier "Shooting Stars" at about 60%. Harvey began the year well ahead of its cohort and still went underwater once agents started working unattended.
A gross margin, for anyone outside finance, is revenue minus the direct cost of delivering the product, as a share of revenue; for an AI application, the model bill sits inside that direct cost. A negative one means every additional unit sold loses money before sales, research or overhead is paid for.
The Fix Was a Model, Not a Price
Harvey's margin recovered on the cost side — it turned positive after Harvey moved work onto Harvey Tenet — rather than through any reported repricing of its seats. Harvey's engineering blog describes Tenet as "a Kimi K3 base that we post-trained together with Fireworks research", and AI Weekly's summary of the reporting describes the result as moving Harvey's flagship product off the frontier labs. The 20 August post states the motive plainly: "open-weight models have cheaper per token prices." On Review Tables, it says the post-trained model delivers better answers "at roughly one-tenth the cost per cell," and on Firm Knowledge it reports "reducing cost per query by 90%."
The token discount alone does not explain that. At list prices today, Fireworks serves Kimi K3 at $3 per million input tokens and $15 per million output, against $5 and $25 for Claude Opus 5 — a 40% gap. A 40% discount does not produce a 90% cost cut. The rest comes from post-training a model to finish tasks in fewer steps, from routing, and — at Harvey's volume — from capacity terms nobody outside the company can see.
Fireworks' own write-up shows the mechanism most clearly. Tenet costs $5.92 per Legal Agent Benchmark task against $5.62 for base Kimi K3, and lifts the all-pass rate from 10.8% to 19.7%. Divide cost by success rate and the cost per fully passed task falls from about $52 to about $30. That — cost per completed task, not cost per token — is the unit your vendor now optimises, and it is the unit you should make them report to you.
Earlier Harvey–Fireworks work showed the routing half. In a June study on a 100-task benchmark slice, open-weight GLM 5.1 did the work and called Claude Opus 4.7 as an "advisor" 0.83 times per task on average. The hybrid fully passed 18 tasks for $368; Opus alone passed 14 for $954. The post's own summary: "The frontier model shows up as a callable tool, not as the dependency the product is built on top of."
The Strongest Case for Harvey's Choice
The honest reading is that Harvey's customers got the least painful of the three fixes. A vendor at -50% gross margin has to change something. Harvey could have metered its lawyers, throttled its heaviest users, or moved work to a cheaper model, and it chose the option that — by its own measurements — raised quality while cutting cost. Investors agreed: Harvey closed a $550 million round at a $15.6 billion valuation this month, with more than 3,000 customer organisations. Nothing in the reporting says the seat price moved — though Artificial Lawyer reported in July that Harvey is understood to be developing a new pricing model for customers, so the meter may be deferred rather than ruled out.
Now the other side. Every quality number in this story is Harvey's or its training partner's, measured on Harvey's own benchmark, and the absolute pass rates are low: the best configuration in either study fully passes fewer than one task in five. "Beats Opus" means 18 against 14 out of 100 — a four-task gap on a sample that size is exactly the kind of difference that often does not survive a confidence interval.
And Harvey has not left the frontier. Anthropic showed investors a slide with a reminder that Harvey still needs Opus for its hardest tasks. So what Harvey's customers now run on is a router. It decides which of your matters are hard enough for the expensive model, and nothing you receive tells you where that line sits — or when it moves.
Harvey does give administrators a lever. Its March post on multi-model design says it lets workspace administrators determine which models are available to their organisation and will disable a provider on request. The same post says Harvey can route work to an alternative model "without disrupting the user's workflow." That capability cuts both ways: the machinery that fails over cleanly during a provider outage is the same machinery that can change your default without anyone noticing.
Every Vendor Has Three Levers. Your Contract Governs One.
The industry has already pulled all three levers, and two of them landed on customers' invoices.
- Meter. Cursor moved its Pro plan to $20 of frontier model usage per month at API pricing in June 2025, explaining that "new models can spend more tokens per request on longer-horizon tasks," then apologised and refunded surprise charges. GitHub followed with usage-based billing for Copilot from 1 June 2026, noting that "a quick chat question and a multi-hour autonomous coding session can cost the user the same amount" — the Harvey problem in one sentence. We tracked what that did to enterprise Copilot bills.
- Cap. Anthropic introduced weekly rate limits from 28 August 2025 after some Claude Code subscribers ran it "continuously in the background, 24/7."
- Swap. Harvey.
None of this is new, only faster. GitHub Copilot was reportedly losing more than $20 a user a month on a $10 subscription in early 2023. It took three years, and the arrival of agents, for that subsidy to reach the price list. Harvey's took one agent update.
A standard SaaS renewal negotiation is built for the first lever: the uplift percentage, the multi-year price cap. It rarely addresses a usage ceiling introduced mid-term as a "fair use" policy, and almost never says which model does the work. The agentic pricing models we compared in August all assume the thing you are buying stays constant while the meter moves. For agentic products, the thing you are buying is now the variable.
Harvey Is the First, Not the Only
The same move is under way in every function where per-seat agent products sell, and in at least one case it predates Harvey. Abridge is building a clinical model on Nvidia's open weights, and Decagon now routes 80% of customer queries through models of its own. Decagon co-founder Jesse Zhang wrote in July that Decagon runs about 90% of its workloads on open-source models, and that latency, not cost, was the real driver.
Ramp raised $750 million in June and is weighing its own model; co-CEO Karim Atiyeh, as quoted by The Next Web, said "It made absolutely no sense a year ago. It's starting to make a lot more sense now." Canva co-founder Cliff Obrecht said the company has to prove profitability, or a path to it, from its AI business, and Canva is shifting to smaller OpenAI and Anthropic models and to open-source ones, while moving to build its own.
Not everyone is persuaded. Menlo Ventures partner Matt Kraning cautioned that building your own model takes specialised staff and higher upfront cost, and told Bloomberg: "In most cases, it tends to be a lot of cosplay." Both can be true at once. Some of these models will carry real production load and some will be a press release — and from the buyer's chair you cannot tell which without asking.
So whether you buy clinical documentation, support automation, finance research or legal AI from Harvey, Legora or anyone else, the model you evaluated at purchase is under the same pressure Harvey's was. The vendors building on Fireworks and similar platforms are not hiding this. Your contract simply never asked.
What to Do Before Your Next Renewal
The work is to get the model, the usage allowance and the unit cost onto paper before the vendor changes any of them for you.
This Week:
- Have your workspace administrator export the list of models enabled in each agentic AI tenant — in Harvey's case, the admin model settings — and date it. That is your baseline for the next change.
- Send each per-seat agentic vendor one written question: for each feature we use, which models served our requests in the last 90 days, and what share went to each? For Harvey customers, the answer shows where the Opus line sits.
- Pull your own usage by user since March. If your top tenth of users runs agents at ten times the median, you are the account your vendor was losing money on — and the account most exposed to a cap.
This Month:
- Re-run your acceptance evaluation against the model serving your tenant today, not the one you tested at purchase. A post-training change can move agent results without anyone's benchmark noticing.
- Draft a model-change clause: 30 days' written notice before a feature's default model changes, the right to pin the previous model for a defined window, and the right to re-run your acceptance tests with a remedy if results fall below what you signed on.
- Get the usage allowance per seat written into the order form as a number. "Unlimited within the seat" can be redefined by a fair-use policy; a number cannot.
Before Renewal:
- Ask for the unit Harvey now optimises: cost or price per completed task. A vendor that reports cutting cost per query by 90% on some workflows is a vendor whose margin just recovered — negotiate for part of that recovery.
- Model both options against your own usage data: a seat with a written allowance, and a metered plan with a hard monthly cap. Choose the one whose worst case you can budget, not the one with the best headline rate.
- Fold model provenance — base model, publisher, licence — into the same clause, so a move to open weights triggers your licence and disclosure review instead of bypassing it.
The Bottom Line
Every flat price on a metered input is a subsidy with an expiry date. Copilot's ran for three years. Harvey's lasted one agent update. The cloud era taught buyers to negotiate the price of compute; the agent era adds a second term — which compute. Harvey has so far kept your price flat by changing what does the work behind it. That may well be the better product. But it is a different product, and it arrived without a contract amendment.
The price held. The model moved. Put both in the contract.
Continue Reading
- CoCounsel's New Model Runs on Qwen. Go Read the Card.
- Agentic AI Pricing: Don't Buy Consumption Without a Cap
- True Cost of a Copilot Seat: The Licence Is the Floor
- Copilot's New Billing Turned a $39 Seat Into $750/Month.
- DeepSeek Swapped the Model. Your Eval Didn't Notice.
- Guardrails' Hub Died Aug 25. Harvey Bought the Team.
