The model that will read your privileged documents inside CoCounsel Legal is a derivative of Alibaba's Qwen, and nothing in your contract records that. Not the DPA. Not the subprocessor list. Not the outside counsel guidelines you sent your firms last quarter. Those documents govern where data flows, and an open-weight base model does not flow anywhere — it is a file the vendor downloaded, modified and now serves from its own infrastructure. The provenance question your GC will ask has no answer artifact.
Two vendors moved within five days. Thomson Reuters launched "Thomson" on 24 August, its first in-house frontier model, describing its base only as "a strong open-source foundation." Harvey published Tenet four days earlier and said plainly that it is "a Kimi K3 base that we post-trained together with Fireworks research." One of those two named the lineage. The other put it in a Hugging Face model card and left it out of the announcement.
What Thomson Actually Is
Thomson is a Qwen derivative, two steps removed, and the only place that chain is written down is the model card. Thomson Reuters published Thomson-1.0-Small on Hugging Face — 35B total parameters, 3B activated, 262,144-token context — and the card states its base checkpoint is "Snowdon1.1-Small." Follow that one more hop and you reach tri-fair-lab/Snowdon1.1-Small, published by the Thomson Reuters–Imperial Frontier AI Research Lab, whose own card names its base model as Qwen/Qwen3.6-35B-A3B with "architecture and tokeniser ... inherited unchanged from Qwen3.6-35B-A3B."
The press release says none of this. CTO Joel Hron gave the specifics to trade press instead, telling Business Insider, as reported by The Next Web, that "there's nothing that necessarily ties us to Qwen." That is true, and it is the most important sentence anyone has said about this launch — because it is also true of every base the company used before this one.
Hron told The Decoder that Thomson Reuters "changed the open source starting point like probably close to a half dozen times already" across more than two years and roughly $40 million in staff and compute, with the final training run costing about $450,000. LawSites reported the most recent foundation as Qwen 3.5; the published card for the small variant says Qwen3.6. Both are Alibaba, and the disagreement is the story: the base moved often enough that careful reporters landed on different answers in the same week.
Thomson's first deployment is inside Tabular Analysis in CoCounsel Legal, where LawSites reports it will be the default model with administrators able to switch. CoCounsel passed one million professionals across 107 countries in February, on a stack that named Anthropic, OpenAI and Google as its model providers. A fourth provider has now joined that list without being named.
Harvey Named Its Base. That Difference Matters.
Harvey disclosed its lineage in its own words, and the honest read is that the disclosure failure here is asymmetric. Harvey's engineering blog says Tenet is post-trained from Kimi K3, the 2.8-trillion-parameter open-weight model Moonshot AI released in July with 104B parameters active per token and a one-million-token context window. Fireworks is named as the research partner. That is more than most vendors volunteer.
It is also not yet in production. Harvey's post presents Tenet as a research preview and says the next step is "scaling our compute so that we can bring our work from research to production" — so nothing about Harvey changed for the 100,000-plus lawyers across 1,300 organizations using it today. Thomson is the one shipping.
The numbers are worth reading carefully. Harvey's own blog reports that Tenet "successfully completes almost twice as many held out tasks on LAB and 20% more on LAB contracts than base Kimi K3, increasing all-pass rate by 9 and 2 percentage points, respectively." The "82% improvement" figure circulating in coverage is that same 9-point gain restated as a relative change off a low base — the underlying result is identical, and Harvey itself publishes only the percentage-point version. On one extraction task Harvey reports a 3.6-point gain in answer quality and 12.1 points in citation quality "at roughly one-tenth the cost per cell." Cost is the stated motive, and Harvey is candid about it: open weights have "cheaper per token prices."
Harvey also states it "did not use any customer data in any of our post-training efforts." Take that at face value. The exposure here is not your documents leaking to Beijing — inference runs on Fireworks' infrastructure in the US, EU and Australia, on weights sitting on disk. The risk is provenance, licensing and disclosure, not exfiltration. Anyone selling you the other story is selling you something.
Your Subprocessor List Cannot Answer This Question
A subprocessor list tracks who touches your data, not whose weights are in the box — which is why open-weight bases are invisible to every procurement artifact you own. Harvey's published subprocessor list, last updated 2 August 2026, names thirteen third-party entities including Microsoft, OpenAI, Google Cloud, AWS, Anthropic, Mistral, DeepL, ElevenLabs, Baseten and Fireworks.ai. Moonshot AI is not on it.
That is not an omission. It is correct. Moonshot processes nothing for Harvey; it published a file. Alibaba processes nothing for Thomson Reuters; it published a file under Apache 2.0 in April. Under GDPR Article 28 there is no further processor to disclose, because no further processing occurs. Every control your firm negotiated — zero data retention, no-training clauses, processing-location commitments, SOC 2 and ISO 42001 attestations — is intact and answers a different question than the one you now need answered.
That is the structural finding, and it generalises well past legal. A DPA is a data-flow document. Model provenance is a supply-chain question, and no document in your procurement file is answering it. When a vendor rented a frontier API, the two happened to coincide: naming the processor named the model. Open weights severed that link. A replacement does exist on paper — SPDX 3.0 has shipped an AI profile since April 2024 that records model lineage, licence and provenance, and CycloneDX publishes an equivalent ML-BOM — but neither vendor here produced one, and an AI bill of materials is not yet something legal procurement thinks to ask for. The Cursor disclosure episode in March was the first version of this problem; the difference now is that the vendors are being reasonably forthcoming and the artifacts customers actually receive still cannot carry the information.
Thomson Reuters' Own Lab Measured the Base Model's Politics
The most substantive work on this launch is in the intermediate model nobody is reading about, and it is both the best evidence of the problem and the strongest mitigation anyone has shipped. The Snowdon card defines "topic-conditioned misalignment" as "the family of responses running from outright refusal, through denial of documented facts, to formulaic deflection, that a model produces on politically sensitive subjects but not on comparable subjects elsewhere." The lab built a re-alignment stage specifically to remove it, drawing on 866 topics relevant to Chinese politics, society and international relations.
The measured results are striking in both directions. On the card's PerspectiveBench re-alignment score across 80 prompts, the model moves from 16.0 to 93.0. Across a larger evaluation of 121,900 prompts spanning 1,219 policies, direct misalignment falls from 44.61% to 14.28%, non-engagement from 16.21% to 2.47%, and policy attack success rate from 54.12% to 21.86%.
Steel-man this properly: that is more transparency about a base model's behavioural failure modes than any enterprise buyer gets from a closed frontier API, and Thomson Reuters published it voluntarily. But read the residual. After remediation, roughly one in seven politically-sensitive prompts still produces direct misalignment, and better than one in five policy attacks still succeeds. That is the number your risk committee will want, and it appears in a Hugging Face card rather than in the launch announcement, the product documentation, or anything a customer receives.
The Licence Runs Backwards
The Chinese base model is more permissively licensed than the Western derivative built on top of it, which is not the direction most procurement teams assume. Qwen3.6-35B-A3B was released under Apache 2.0 — use it, change it, redistribute it, sell it. Thomson-1.0-Small ships under PolyForm Strict 1.0.0, which grants rights for permitted purposes "other than distributing the software or making changes or new works based on the software," and limits use to noncommercial purposes: personal study, charities, educational institutions, public research organisations, government. LawSites characterised it as a non-commercial academic licence, and that is right.
So the "open-weight" release is a look-but-don't-touch artifact. If your firm was planning to pull the weights and evaluate them against your own matter set, read the licence before your engineers download anything — an internal evaluation supporting a commercial purchasing decision is not obviously a permitted purpose. Thomson Reuters has said a technical report and a developer portal are to follow, which is where any real customer use will live.
Kimi K3 is its own puzzle. The Kimi K3 licence requires a separate agreement with Moonshot AI before any commercial use once "the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars ... in total over any consecutive 12 months" — note that the threshold is measured on your revenue, not on what the model earns — and mandates that "Kimi K3" be prominently displayed in the user interface of products above 100 million monthly active users or $20 million in monthly revenue. Both obligations fall away for purely internal use, and for access through Moonshot's own products or its "certified inference partners," which is the carve-out to check before assuming a threshold bites. Whether Harvey crosses either line is a question for its lawyers, not yours. But if you are running open weights yourself, those thresholds are exactly the clauses that turn a free download into a negotiation two years later — the same class of exposure covered in the Hugging Face weight-mirroring playbook.
What the Benchmark Table Does Not Compare
Thomson Reuters claims parity with "the latest frontier models," and its published comparison set contains no frontier model. The card benchmarks Thomson-1.0-Small at an overall average of 74.6 against Snowdon-1.1-Small at 71.7, Qwen3.6-35B-A3B at 71.7, Gemma 4-31B at 71.2 and Haiku 4.5 at 68.2. Haiku is Anthropic's small, cheap tier. There is no Opus, no Sonnet, no GPT-5.x in the table.
On the legal domain average — the whole point of the exercise — Thomson scores 75.2 against the raw Qwen base at 72.7. Two and a half points, for $40 million and two years of Westlaw, Practical Law, Checkpoint and Reuters content on the company's own benchmark. Note also that Snowdon-1.1-Small scores 72.4 on legal, marginally below the unmodified Qwen it was built from: the political re-alignment stage appears to have cost a little raw capability, which is what safety work usually does and which nobody claimed otherwise.
The Decoder reports the more useful number. On Thomson Reuters' internal Deep Research benchmark with web access, Thomson scored 0.53 on factual accuracy against GPT-5.4's 0.65, and only pulled ahead — 0.83 to 0.82 — once it had access to the company's proprietary content. That is a completely defensible result for a 3B-active model, and it is a precise statement of where the thing is good: inside the walled garden, on the corpus it was specialised for. It is not a frontier model and the benchmark table does not claim to be one. The press release does.
The Rules Are Entity-Specific. Derivatives Are Not Covered.
Every restriction currently on the books names a company, and none of them reaches a model post-trained in London from weights published in Hangzhou. A bipartisan bill from Senators Bill Cassidy and Jacky Rosen would bar federal contractors from using DeepSeek for any activity related to a federal contract, extending to successor models developed by High-Flyer. The No Adversarial AI Act, sponsored by Senators Rick Scott and Gary Peters with a bipartisan House group, would have the Federal Acquisition Security Council identify and publish a list of AI developed by companies associated with China, Russia, Iran and North Korea. Neither addresses a derivative.
Lists name entities. A Qwen derivative, re-aligned by a UK university lab, continually pre-trained on American case law and served from Thomson Reuters infrastructure is not on any list and would not be easy to put on one. If your firm has federal, defence or export-sensitive clients whose engagement letters carry model-origin clauses, that ambiguity is yours to resolve, not the statute's. The Pentagon's exclusion of Anthropic over supply-chain concerns showed how fast a procurement position can move on origin grounds.
One more live thread. Anthropic wrote to US senators and White House officials on 10 June alleging that operators linked to Alibaba's Qwen team used nearly 25,000 fraudulent accounts to run some 29 million exchanges against Claude between April and June, targeting software engineering and agentic reasoning. Alibaba had no comment; the allegation is unproven and has not been tested anywhere. Treat it as exactly that. But if it were ever substantiated, the IP-indemnity question would reach downstream into every derivative — which is precisely the sort of contingency your indemnity clause was written for and almost certainly does not name. We covered the frontier labs' earlier distillation-intelligence pact when this first surfaced.
What to Do About It
This Week:
- Open your CoCounsel Legal admin settings before the next release lands and find out whether Tabular Analysis is defaulting to Thomson in your tenant. LawSites reports administrators can switch. Decide deliberately rather than by default.
- Send both vendors one question in writing: for each model available in our tenant, name the base model, its publisher and its licence. Not "do you use Chinese models" — that invites a defensible no. Ask for lineage.
- Read Harvey's subprocessor list yourself and confirm what it does and does not tell you. It is the artifact your procurement team believes answers this question, and it does not.
This Month:
- Add a model-provenance clause to your AI vendor questionnaire: base model, publisher, licence, and written notice when any of the three changes. Thomson Reuters changed its base close to six times while building this — during development, so no customer was owed notice — and nothing in a standard DPA would oblige a vendor to tell you about the next swap either. That gap is a contract term, not a scandal.
- Run your top three legal-AI vendors through the same question and see how many can answer within a week. The ones that can are the ones that have thought about it.
- Check whether your outside counsel guidelines and client engagement letters define "AI provider" in a way that survives an open-weight base — most define it as a processor, which no longer captures the model's origin.
Before Renewal:
- Rebuild your evaluation set against whichever model is actually serving your tenant. A silent base swap invalidates the eval you ran at purchase, exactly as it did when DeepSeek changed its post-training and nobody's benchmarks noticed.
- Escalate model-origin questions to your federal and defence-sector clients now, before an engagement letter forces the conversation on someone else's timetable.
- Price the switching cost. If a provenance answer would fail your client's requirements, you need to know today what it costs to move off the feature, not during a renewal negotiation.
The Bottom Line
Neither of these vendors did anything wrong. Harvey named its base model in its own engineering blog. Thomson Reuters published a benchmark table, a licence, and an unusually candid measurement of its base model's political behaviour — more disclosure than any closed API vendor offers. The problem is that both did it in places procurement does not read, in a week when the 46% of enterprise AI traffic already running on Chinese models quietly gained a category nobody was counting: Western vendors' own flagship models, built from Chinese weights, served from Western infrastructure, invisible to every control in the contract.
The legal AI market grew up promising that the model touching privileged work would be named, contracted and auditable. Harvey on Claude was that arrangement. It is being replaced by something cheaper and better-performing, and the paperwork has not caught up.
The model card is now more informative than the contract. Read the card.
Continue Reading
- Cursor's Hidden Chinese Model: A Wake-Up Call for AI Vendor Due Diligence
- 46% of Your AI Now Runs on Chinese Models
- Hugging Face Hired Bankers. Go Mirror Your Weights.
- DeepSeek Swapped the Model. Your Eval Didn't Notice.
- Pentagon Blocks Anthropic: Chinese Chips Disqualify AI Leader
- Anthropic's Dreaming Made Harvey's Agents 6x Smarter
- 16M Stolen Queries Force AI Labs to Share Intel
- $1.2B Legal AI: Why Blackstone Pays for Outcomes, Not Hours
