Mistral Large 4 Beats Kimi K3 on Legal Work With No Licence Yet

Mistral Large 4 beats Kimi K3 on Harvey's legal agent benchmark and trails it on the Vals Index. It is a discounted preview API for now, and the licence for its end-of-October weights is unpublished.

By Rajesh Beri·October 6, 2026·9 min read
Share:
An eight-GPU server sled being slid halfway into an otherwise empty rack in a cold datacenter aisle, with a thick unsigned paper contract resting on top of the rack, its signature page blank.

Illustration generated using AI

Mistral Large 4 is worth a benchmark run this week and not yet worth a sovereignty promise. On independent tests it beats Kimi K3 on Harvey's legal agent benchmark and on Terminal-Bench 4, but it trails Kimi K3 and GLM-5.3 on the broad Vals Index. Today it exists only as a preview API at a discounted launch price. The weights are due by the end of October under a licence Mistral has not published.

Mistral announced Large 4 in public preview on October 6, calling it a 1 trillion parameter mixture-of-experts model with 49 billion active parameters, text and image input, and training in more than 160 languages. The nickname is "Le Chonk." For an EU bank or insurer currently choosing between a Chinese open-weight model, a US one and a hosted frontier API, it is a European entrant at this size. Whether it belongs on your shortlist comes down to three questions the launch post doesn't fully answer: how it scores when someone else runs the test, what it costs per finished task, and what you will be allowed to do with the weights.

Where Mistral Large 4 Beats Kimi K3, and Where It Doesn't

The independent numbers put Large 4 ahead of the open-weight field on legal agent work and behind it on general agentic work. Vals AI published results the same day, and they are more mixed than Mistral's own charts.

On Harvey's Legal Agent Benchmark as run by Vals, Large 4 completes 15.83% of tasks at $4.18 per test. Kimi K3 completes 12.92% at $3.83 and GLM 5.3 completes 8.33% at $4.30. Gemini 4 Argon scores 19.58% but costs $10.55 a test, and Claude Opus 5.5 lands at 3.75% for $21.38. Those pass rates look low because the metric counts a task only when every rubric criterion passes. Vals notes that top models satisfy around 90% of individual criteria. For a legal team, that gap between the two numbers is the review burden you would still carry. If legal drafting is your use case, this is the strongest result in the launch, and it matters because Harvey itself has been moving work onto open-weight models to protect its margin.

On Vals' Terminal-Bench 4 run, Large 4 scores 22.73% against Kimi K3's 17.17%. GLM 5.3 scores 38.89%, and Claude Opus 5.5 leads at 65.15%. Large 4 is a usable coding model and nowhere near the hosted leaders.

The composite scores reverse the legal result. The Vals Index v2.1, updated October 6 and weighted across finance, coding, legal and tax work, puts Large 4 at 48.05%. Kimi K3 scores 50.30% and GLM 5.3 scores 53.51%. On the Artificial Analysis Intelligence Index, the preview scores 38. Kimi K3 (Max) scores 44.

Mistral's own coding chart, reported by VentureBeat, shows Large 4 at 62% on DeepSWE v1.1 against 61% for GLM-5.3 and 44% for Reflection's Beam. VentureBeat points out that the public DeepSWE leaderboard shows GLM-5.3 near 69% under a different configuration, and that several competitor figures in Mistral's charts could not be traced to a public source. Treat the vendor charts as a claim and the Vals pages as the evidence. Our read of Reflection's Beam launch found the same pattern: a new open-weight model wins a narrow slice and loses the composite.

Half the List Price, Twice Kimi's Cost per Vals Task

Large 4 is cheaper than Kimi K3 per token and, on the Vals Index, more expensive per task. That difference comes from how much the model writes.

List pricing is $1.36 per million input tokens and $4.18 per million output tokens. Mistral's model page currently shows $0.68 input, $0.07 cached input and $2.09 output, with the $1.36 list price crossed out. That is a launch discount. Budget on the list price, because the discount is what you will lose first. Kimi K3 (Max) is listed by Artificial Analysis at $3.00 input and $15.00 output, so on paper Mistral costs less than a third as much per output token.

Per task, the picture changes. Artificial Analysis found the Large 4 preview "very verbose": it generated 200 million output tokens during its index run against a median of 81 million. On the Vals Index, Large 4 costs $13.78 per test against $6.38 for Kimi K3. On Artificial Analysis's own index, it still comes in cheaper per task ($1.13 against Kimi's $2.00). Two evaluators running different harnesses give opposite answers on cost. Your own workload, with your prompts and your reasoning settings, is the only number that will hold in a budget review.

The context window needs the same check. Mistral's model page says 1 million tokens. Artificial Analysis lists 524,000. If your plan depends on stuffing whole contract sets into one call, test the long end before you design around it.


The Licence Isn't Published, and Mistral Has Shipped Three Kinds

"Open weight" is a promise until the licence file exists. VentureBeat reports the weights will ship on October 27 under a custom Mistral licence. Mistral's own post says only "end of October," and the terms are not public.

Mistral's recent history gives you no default to assume. Mistral Large 3 shipped in December 2025 under Apache 2.0, with no field-of-use or revenue strings. Medium 3.5 came under a modified MIT licence with exceptions for companies with large revenue. A custom licence for the flagship is a third path, and custom open-weight licences are where revenue thresholds, attribution rules, acceptable-use lists and restrictions on serving the model to third parties tend to live. Until you read it, you don't know whether a bank can fine-tune it, whether a systems integrator can host it for clients, or whether a revenue cap applies.

This is the same diligence we recommended when Mistral raised its Series D: score the contract clauses, because the country on the letterhead tells you little about what you can do with the model. A European lab with a restrictive licence gives you less control than a Chinese model under a permissive one that you host yourself. Kimi K3's own licence allows commercial use with restrictions, so neither shortlisted option is Apache-clean today.

There is also a sovereignty point inside the preview. Mistral says it trained the model on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters, but the API will be served from "multiple regions worldwide," one of them a European deployment Mistral runs end to end. When we looked at Mistral's EU regional endpoint in August, it dropped agent features that the global endpoint carried. Confirm which region your preview traffic runs in before a compliance team signs off on the pilot.

What It Takes to Self-Host 1.05 Trillion Parameters

Plan for one eight-GPU Blackwell node at 8-bit precision and two at 16-bit, and treat both as an estimate until a supported checkpoint ships. Mistral's model page gives the size more precisely than the launch post: 1.05 trillion total parameters and 52 billion active, with a 1.6 billion parameter vision encoder.

Active parameters set the compute per token. Total parameters set the memory bill, because every expert has to be resident on the GPUs. We worked through this when Tencent's 49B-active model still needed room for all 770B. The arithmetic for Large 4 is simple: about 1.05 TB of weights at one byte per parameter, about 2.1 TB at two bytes, and roughly 525 GB at four bits.

An NVIDIA DGX B200 has 1,440 GB of GPU memory across eight GPUs. At 8-bit, Large 4 fits in one node with roughly 390 GB left for KV cache and runtime overhead. That headroom is thin for million-token contexts and many concurrent users. At 16-bit you need two nodes and a working multi-node serving setup. Kimi K3 is 2.8 trillion total and 104 billion active, so by the same arithmetic it needs more than two and a half times the memory. If your hosting decision was going to be Kimi K3 on two nodes, Large 4 could cut that to one. That is a real advantage, and it depends on an FP8 checkpoint and runtime support in vLLM or SGLang that do not exist yet.

None of this is validated by Mistral. Until the checkpoint ships, no one outside Mistral can tell you the real memory footprint, the quantisation loss, or the throughput on your hardware.


What to Do Before the Weights Land

This Week:

  1. Run your own eval set against the preview API on the discounted price. Pick 50 to 200 tasks from the function you care about most (contract review, credit memos, ticket triage), and run the same set against Kimi K3 and your current hosted model. Record pass rate, output tokens per task and cost per task.
  2. Log which region your preview calls are served from and put the answer in writing before any regulated data touches it.
  3. Price the pilot at the $1.36 / $4.18 list rate in your business case, and treat the launch discount as a bonus.

Before October 31:

  1. Get the licence text the day it lands and send it to legal with four questions: is there a revenue or user threshold, can you fine-tune and keep the derivative private, can a third party host it for you, and is there an acceptable-use list that covers your industry.
  2. Mirror the weights into your own registry as soon as they publish, with a hash, so your pilot cannot be broken by a later re-upload. Our weight-mirroring playbook has the steps.

Before You Commit Q1 GPU Capacity:

  1. Ask your infrastructure team to size one eight-GPU Blackwell node for an 8-bit checkpoint and two for 16-bit, then hold the purchase until a quantised checkpoint and a supported runtime exist.
  2. Rerun your eval on the self-hosted checkpoint. The preview API and the published weights are not guaranteed to be the same model, and Mistral's chief scientist Guillaume Lample told VentureBeat the company expects capabilities to improve as reinforcement learning concludes.

The Bottom Line

Large 4 changes the European option from a mid-sized model to a frontier-sized one, and on legal agent work it now beats the Chinese model most buyers were going to use instead. It has not closed the gap on composite agentic scores. Its per-task cost depends on whose harness you believe, and the most important document in the decision, the licence, does not exist yet.

The last time a European lab shipped an open flagship, Large 3 came under Apache 2.0 and buyers could plan around it on day one. This time, keep the commitment small until October 31. Benchmark on the API now, read the licence when it publishes, and size the hardware after the checkpoint is real.

Continue Reading

Share:

Frequently Asked Questions

How much does Mistral Large 4 cost?

List pricing is $1.36 per million input tokens and $4.18 per million output tokens. During the preview, Mistral's model page shows a launch discount of $0.68 input, $0.07 cached input and $2.09 output. Budget on the list price, and note the model is verbose: on the Vals Index it cost $13.78 per test against $6.38 for Kimi K3.

Is Mistral Large 4 better than Kimi K3?

On legal agent work, yes: Vals ran Harvey's Legal Agent Benchmark and Large 4 passed 15.83% of tasks against 12.92% for Kimi K3. It also leads on Terminal-Bench 4. On composite scores it trails: 48.05% against 50.30% on the Vals Index, and 38 against 44 on the Artificial Analysis Intelligence Index.

When will Mistral Large 4 weights be released, and under what licence?

Mistral says the weights will ship by the end of October 2026; VentureBeat reported October 27. They are expected under a custom Mistral licence whose terms have not been published. Mistral Large 3 used Apache 2.0, so do not assume the same terms.

What hardware do you need to self-host Mistral Large 4?

Mistral lists 1.05 trillion total and 52 billion active parameters, so all experts must sit in GPU memory: about 1.05 TB at 8-bit and 2.1 TB at 16-bit. One DGX B200 with 1,440 GB of GPU memory fits an 8-bit checkpoint with limited KV cache headroom; 16-bit needs two nodes. These are estimates until a supported checkpoint ships.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →