Skip to main content

2 posts tagged with "vendors"

View All Tags

Is Post-Training Enough to Make Chinese Base Models OK for Law? Thomson-1 (Qwen) and Harvey Tenet (Kimi) Are Some of the Riskiest Positioning of Chinese Models

· 9 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC

DOJ investigations. eDiscovery in class action lawsuits involving critical intellectual property in high-tech industries. A case involving transnational money laundering or a foreign student taking photos near airports or military bases. Counterintelligence investigations. Embarrassing kompromat that is supposed to be protected by privilege or other legal doctrines. Myriad other private details and secrets that could be leaked to and compiled by Chinese intelligence.

Do I know for certain that this’ll happen? No. But, you also cannot guarantee that a model created by an adversary has no backdoor, it is extremely naive to the point of absurdity to pretend that the adversary (the PRC) is actually not a threat or that legal workflows are not an incredibly valuable target for espionage, and backdoor triggers can be very obscure and seemingly benign (see, e.g., “Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs”).

The NSA, CISA, and FBI jointly released an advisory yesterday, September 8, 2026, China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies, AA26-251A (U/OO/6059854-26 | PP-26-3853 | September 2026 Ver 1.0), listing specific capabilities that get distilled from American AI models to Chinese models. The first listed use case examples of DeepSeek distillation include law:

Between late 2024 and mid-2025, DeepSeek [DeepSeek (DeepSeek Artificial Intelligence Technology Research Co., Ltd.) 深度求索AI基础技术研究有限公司] distilled specialized training data and capabilities from the following U.S. frontier AI company [Anthropic, OpenAI, Google, and xAI] models to train their R1 and V3 models: [] The specific knowledge and capabilities distilled included:

  • Legal specialization optimization
  • API rule-driven tasks
  • Writing using CoT drafts
  • Agentic functions
  • Question and answer optimization
  • Coach/assistant capabilities
  • Functional creation optimization
  • Supervised fine-tuning (SFT) optimization
  • Creative and occupational writing optimization

The joint advisory goes on to accuse “Moonshot AI (Beijing Moonshot Technology Co., Ltd.) 北京揽月星辰科技有限公司” of distilling from a variety of OpenAI, Anthropic, Google, and xAI models, including Claude Fable 5, GPT-5 Codex, Nano Banana (the highly capable Gemini image generation model), and xAI Grok Code Fast-1. The advisory does not explicitly mention legal specialization for Moonshot’s Kimi-K2.

  • Harvey’s own research preview, published August 20, 2026, states: “Harvey Tenet is a Kimi K3 base that we post-trained together with Fireworks research for long-horizon legal work.” NOTE: This appears to be a reference to the U.S.-based company Fireworks.AI, which offers post-training on a variety of Chinese base models, as well as Nvidia’s Nemotron.

The advisory mentions “Alibaba 阿里集团,” maker of the Qwen family of models, distilled from Anthropic and OpenAI models for tasks including “[e]nd-to-end agentic workflows,” which would be helpful in more automated legal workflows.

  • Thomson-1 is built on Alibaba’s Qwen models, according to reporting from Business Insider; Thomson Reuters’ technical report, Thomson: Continual Learning of Frontier Models for SovereignAI, states that Thomson-1.0-Large and Thomson-1.0-Small started from Qwen3.5-397B and Qwen3.6-35B respectively. An intermediate “value-realigned” version, called “Snowdon,” was produced with Imperial College London before the legal post-training to create the “Thomson” models. The August 20 press release calls the model simply “Thomson” and does not mention Qwen or China.

Thomson-1 and Harvey Tenet: What "Post-Trained" Is Doing In That Sentence

Two of the largest legal AI vendors both announced in late August 2026 that they were releasing a flagship model built on a Chinese open-weight base model with post-training on top, in a likely bid to manage token costs from one or more of Anthropic's Claude, OpenAI's GPT, or Google’s Gemini frontier model APIs.

  • Thomson Reuters: Westlaw CoCounsel runs on a mix of Claude models. TR announced the new “Thomson” or “Thomson-1," described as post-trained, with TR’s own technical report (but not its press release) acknowledging Alibaba’s Qwen open-weight models as the base.
  • Harvey: Harvey has a mixture of model options, with the model router by default selecting from a number of Claude, Gemini, and GPT models. "Harvey Tenet," per Harvey, is built by post-training a Chinese open-weight model, Kimi K3. According to Simon Willison’s analysis of the Kimi K3 license, a “Model as a Service business” would require a separate agreement with Moonshot if “the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars [] in total over any consecutive 12 months...”

Damien Charlotin (of Hallucination Database Fame) Makes Great Point About Being Lulled Into a False Sense of Security by Legal AI

· 4 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC

“What 1200+ Hallucinated Citations Teach Us About Legal AI,” with Rok Popov Ledinski. https://www.youtube.com/watch?v=K20Kprb7cbo&t=980s

Key quote for me came at roughly 16:20-18:00 [transcribed by me and edited slightly for clarity]:

More generally, a kind of new pattern that you have is that you include the hallucinations in your brief. You trust AI to write legal content, because it works so well in a lot of other contexts. And it even works so well in the legal field. So if you're using a specialized tool, if you're using-especially if you're using one of the typical traditional platforms you have been using for years. You have been a loyal client to Westlaw for years. They have always been working very well to give you the legal content that you wanted. They've got an AI tool. Why would you not trust this AI tool? Especially when the marketing team of—and I don't want to target Westlaw or anything, that could be any other legal editor—"well, we don't hallucinate because we've got all these kind of methods to make sure we don't hallucinate." Then it's not a question of being aware of hallucination. It's just: you're trusting a tool that tells you that there's no hallucination. You're trusting it because they're telling you, but also because you've got your own habits of them working. And sometimes, if the tool is actually good, 99% of the cases—of the times—it will not hallucinate and you're fine with it. And, you know, when people like me come around and say "You should probably try to check it every time," but if you check 99 times and nothing is wrong, at some point you'll stop checking. And that's completely human and completely normal. So I'm not sure the answer to this is "continue checking just in case." I think the answer will be partly technological and partly still a bit of checking and layering and stuff like that.

A Few Quick Observations

  • I appreciate this point, because I’ve been frustrated by the lack of clarity from GenAI software vendors on basic risks like hallucinations and prompt injection, whether they're making business dashboards, or email and scheduling tools, or legal research tools. Across the industry, downplaying these risks has led users to accept the “models are getting better” narrative.
  • What Charlotin describes here is basically a “normalization of deviance” (a phrase popularized after the Challenger Disaster). Every time someone skips a verification step and something bad does not happen, it helps rationalize not checking for errors. But even if we accept lower error rates, stuff will still get through if we don’t check at all. But lower error rates might make us complacent.
  • Commentators on generative AI in fact frequently invoke the Challenger comparison explicitly, e.g., Simon Willison’s 2026 predictions for coding agents.
  • This is addressed by my “glass donut” metaphor, described in my previous post about Sullivan & Cromwell. If it is true that AI hallucinations are getting less frequent, but nonzero and still catastrophic when they occur, the practical effect may be to lower our guard and cause bad habits that fail eventually. Charlotin argues that technology probably has to be part of the solution, and indeed I’m experimenting with a word processor to address some of these core problems as a side project. Charlotin’s PelAIkan cite checker is also an approach. But he says and I agree that training and checking will still be a part of the mix.