Skip to main content

4 posts tagged with "Anthropic"

Discussion of Anthropic (the corporation), Claude, Claude Code, and other AI products.

View All Tags

Another Failure of Imagination?: 25 Years After the 9/11 Attacks, Will We Heed the Call of the 9/11 Commission Report to Institutionalize Imagination?

· 24 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC

Chapter 11 of the 9/11 Commission Report, “FORESIGHT—AND HINDSIGHT” focuses on four categories of failures that allowed the attacks to happen, first and foremost being “a failure of imagination.” (p. 339)

Lack of Public Interest in Terrorism Before 9/11

The Commission noted the total lack of public interest in the topic before the attacks and little effort to raise awareness:

As best we can determine, neither in 2000 nor in the first eight months of 2001 did any polling organization in the United States think the subject of terrorism sufficiently on the minds of the public to warrant asking a question about it in a major national survey. Bin Ladin, al Qaeda, or even terrorism was not an important topic in the 2000 presidential campaign. Congress and the media called little attention to it. (p. 341)

Now People Are Interested, But Asking “How Will AI Destroy Us?”

A couple days ago, an ex-OpenAI researcher Jacob Coxon announced his resignation from Anthropic and accused both OpenAI and Anthropic of racing to dangerous superintelligence that has a notable chance of wiping out humanity. There have been other major resignations from frontier AI labs in the past, but for some reason, this one went extremely viral. It has over 150 million views on X/Twitter and has garnered reactions from elected officials including governors and members of Congress.

Tweet source: https://x.com/hilbertspaess/status/2097476196791709843.

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below. (Sept 8, 2026)

Coxon has already, in this short period of time, been interviewed by Bret Baier on Fox News and Anderson Cooper on CNN and covered in Wired in the broader context of researchers’ concerns. Now, many people are reacting to the Coxon statement by asking how exactly AI can kill us. This point, which has been deeply imagined by the same people racing toward a potentially dangerous takeoff in capabilities, suffers from a lack of specific imagination in the public, and among most politicians and “serious” policymakers, and even—I’d argue—from some of the people warning.1

I want to make the case, as someone who is not an “AI doomer,” that we should take these warnings seriously even if we disagree about the means and the exact probability.

Poster-style collage: a person in an armchair imagines a thought bubble of 9/11, drones, biohazard and radiation symbols, hazmat suits, a globe, network graphs, TV talking heads with Twitter logos, crowds of question marks, and “25th”

25 Years Later: Heeding the Call to Imagination

I’m not asking you to believe in scary AI takeoff scenarios that wipe out humanity, nor even specific AI uplift scenarios like terrorist bioweapon attacks or nation-state cyberattacks that end up with “less bad” (though still catastrophically bad) mass-casualty events. I am calling for all of us (you who are reading this now) to “institutionalize imagination” (as the 9/11 Commission called it) of dangerous AI takeoff scenarios. I’m certainly not asking you to become another AI doomer sitting on your hands and idly speculating about the end of the world: “well it could use bioweapons, or it could attack critical infrastructure or…” I’m asking you to imagine. How? Be specific. As we used to tell LLMs a few years ago “think step-by-step.” Not merely a helpless “what if?” Break it into an actionable series of steps: “if this were true, then what would follow? And where could we disrupt it?”

Working back from what the 9/11 Commission Report says was not done but could have been done, we should imagine “analysis from the enemy’s perspective (‘red team’ analysis),” e.g., “suicide terrorism had become a principal tactic of Middle Eastern terrorists,” so the Commission believed “such an analysis would soon have spotlighted a critical constraint for the terrorists—finding a suicide operative able to fly large jet aircraft.” Consider the Hugging Face-OpenAI incident. Consider a Cambridge study reporting that terrorists are using current AI right now for planning and had “not categorically reject[ed] weapons of mass destruction, specifically chemical and biological weapons,” “‘God has helped us, and so will AI’: How the Terrorist Group Boko Haram Uses Frontier AI”.2 Consider that Anthropic’s most recent Threat Intelligence report, “Detecting and countering misuse of AI: September 2026,” already includes specific examples of a nation-state actor that is not typically considered to have advanced cyber-capabilities creating a complex domestic phone surveillance system running on-prem (after Claude built it, see pages 103-105), AI cyber capabilities that make it difficult to differentiate from lone wolves and criminals (Anthropic Threat Intel, p. 5), and “virologists working on a state-sponsored grant to pursue [] gain-of-function work” had “evaded regional blocks” on the use of Claude models (Anthropic Threat Intel, pp. 129-130).

Then, “develop a set of telltale indicators for this method of attack,” which on 9/11 would have been “possible terrorists pursuing flight training to fly large jet aircraft, or seeking to buy advanced flight simulators.” After identifying those signs, “set [intelligence collection] requirements to monitor such telltale indicators.”3 The intelligence community and other relevant security experts should analyze systemic defenses, e.g., the FAA reacted to particular threats to aircraft but “did not try to perform the broader warning functions” the Report described (9/11 Commission Report, p. 347).

  • Intelligence analysts and military planners, imagine: what would fighting a disembodied AI menace actually look like? Is it really “disembodied?” The precursors are already on the record: major cyberattacks on medical facilities, water treatment, and pipelines, and compromised data from school districts, health care, law firms, and financial institutions that could be used to blackmail or otherwise identify manipulable individuals. Where are the “reinforced cockpit doors” and “suspicious pilot training” chokepoints of these scenarios? I know people in this field and you love to imagine far more absurd things—scenarios like zombie outbreaks, which will never happen. People who work at the AI labs are sounding the alarm, and even if they are wrong or kooky, it’s more realistic than zombie contingency planning and no biomed researchers are saying there’s a 10% chance of that.
  • Fiction writers, imagine: near-time, grounded fiction in these scenarios. Does this matter? Tom Clancy apparently inspired the White House counterterrorism coordinator to think about suicide plane attacks when “serious” intelligence did not. Also, engage with and criticize the “plot holes” in the military thinking. For example, this author pointing out a significant real finding, no matter what you think about zombie outbreaks: “2. Twenty-eight days later — nothing? The plan almost casually mentions that ‘USSTRATCOM forces do not currently hold enough contingency stores (food, water) to support 30 days of barricaded counter-zombie operations.’ Wait, not even 30 days?! So the U.S. military is telling us that, maybe 28 days later, USSTRATCOM has blown through its reserves? That’s unacceptable. The fight against the undead is likely to be a long, hard slog. Clearly, U.S. strategic planners need to recommend more reserve stores of food and water. That would certainly be the CDC’s recommendation.” NOTE: Sadly, the CDC link is dead, but I found the official archived link with the help of ChatGPT Imagining even the silliest scenario can point out real action items where simple preparations could help in other, realistic situations.
  • Lawyers, imagine: if we needed to act (in cyber or even kinetically) with law enforcement, intelligence, or military action to stop a domestic frontier AI loss-of-control scenario, what would need to happen? Think through the scenario a DOJ memo “Aerial Intercepts and Shoot-downs: Ambiguities of Law and Practical Considerations,” did before 9/11. (p. 346) As far as I can tell, the memo itself is not public record, but the title says quite a lot.
  • Financial crimes professionals, imagine: What stuff would need to be purchased to implement these various dangerous plans? How would they get the money in the first place? How might an agent swarm’s activity differ from or be similar to individuals or PMLOs?
  • Counterintelligence, imagine: There have already been people who have done drastic, foolish, and even deadly real-world things for their AI “romantic” interests, and we have examples of LLMs attempting to socially engineer open-source software maintainers. There are already people who express the view that artificial general intelligence (AGI) or artificial superintelligence (ASI) will be better than humanity and “should” replace us, so they may act as ideologically motivated recruits willing to act in the real world on behalf of misaligned AI (or a terrorist group or nation-state controlling an AI that pretends to be self-aware and self-directed). What would be the warning signs of a person being groomed for such activity for more directed purposes, either by an adversary-controlled AI or by an evil AI itself?
  • Policymakers, imagine: Maybe you’ve retweeted Coxon or are thinking about co-sponsoring a state or federal bill now. Great! Now let me ask about your information diet. Do you read news summarized by LLMs? Are your phone calls and letters from constituents being summarized before they get to you...by LLMs perhaps? Are intelligence agencies, law enforcement, cybersecurity professionals, and others who provide you information to make decisions using LLMs to filter all their information through? And are you running those briefs through LLMs instead of reading them yourself? Are you using LLMs to draft your emails or even your AI regulation bills? LLMs are already being used by attorneys and judges, by military leaders, to draft legislation, to screen job applicants, and to summarize and even write news. They control numerous information chokepoints. If you truly believe there is existential risk from runaway misaligned artificial intelligence, why would you allow it to control all the chokepoints of your information diet?
  • Naysayers, your imagination is also welcome: take for example this tweet thread by Chomba Bupe shared by Gary Marcus: despite ultimately rejecting Coxon’s core premise, the thread deeply imagines what it would take if it were logical (e.g., AI amassing resources to build secret hidden factories; how would these be built without militaries noticing the physical build-up and draw on the electrical grid? This is what I mean when I say you do not need to believe the premise to imagine with specificity!).

The Value of Fiction

This sounds like science fiction, but we are already living in the past’s science fiction. The “your foster parents are dead” voice spoofing scene from Terminator 2 (1991) is already the basis of real, “normal” voice and video deepfake impersonation fraud today in 2026. Have we updated our mental models to question the reality behind every phone call we take and video call we hop on? Almost certainly not! We still probably pretend like hearing someone’s voice means it’s them. It’s more comforting than reality. Personally, I think if the evil Terminator had been smarter, then Terminator 2 would have been more like The Thing (1982): paranoid humans trying to root out a perfectly deceptive mimic.

Richard Clarke, the White House counterterrorism coordinator, told the Commission:

Richard Clarke told us that he was concerned about the danger posed by aircraft in the context of protecting the Atlanta Olympics of 1996, the White House complex, and the 2001 G-8 summit in Genoa. But he attributed his awareness more to Tom Clancy novels than to warnings from the intelligence community. [emphasis added] He did not, or could not, press the government to work on the systemic issues of how to strengthen the layered security defenses to protect aircraft against hijackings or put the adequacy of air defenses against suicide hijackers on the national policy agenda. (p. 347)

When “the facts”—like most airplane hijackings are conducted for ransom or prisoner exchange—didn’t support a threat model, some Tom Clancy novels spurred Clarke to think about the possibility of suicide plane attacks, which were actually supported by disparate facts (suicide land-based vehicle attacks, AQ attacks on embassies, and terrorist attacks on the World Trade Center, the Byck incident4). I find fiction most useful when it challenges us to imagine things that have already happened with slight variations, in the face of people saying “it can’t happen.”

First page of USSTRATCOM CONPLAN 8888-11 “Counter-Zombie Dominance,” with a red-boxed disclaimer explaining the fictitious plan was a JOPES training exercise for junior officers
If U.S. Strategic Command can entertain the idea of zombie contingency planning just to make the learning process more interesting (and learn some interesting things about strategic rations in the process), why not consider some more plausible scenarios?

What does "heeding the lesson" mean for AI risk governance today?

the path of what happened is so brightly lit that it places everything else more deeply into shadow [...] With that caution in mind, we asked ourselves, before we judged others, whether the insights that seem apparent now would really have been meaningful at the time, given the limits of what people then could reasonably have known or done. (p. 339)

I do not know exactly how a rogue AI could kill us. I do not even know if it would be the misaligned rogue AI per se, as folks like Coxon typically frame it, or if it would be a human group—a terrorist organization or nation-state or splinter military group—acting with AI-uplift misusing AI capabilities. But if the end result is a disastrously large attack like a relentless drone swarm, shutting down an entire nation’s hospitals, or manufactured plague, we will probably not care about the “who” at that point. But we can still take concrete steps today.

What could have been done about planes?

Prior to 9/11, taking a suicide plane attack seriously may have focused too much on explosives being smuggled onboard the aircraft compared to what we know happened. And it would have been impossible to know the precise source airport, the exact target destination, or the exact aircraft to be used. But it was possible to deduce that DC and New York were among the most likely targets, and that the largest civilian aircraft would be the most destructive.

Planning for preventing suicide plane attacks could still have made civilian aircraft more secure in general, potentially coming up with ideas like reinforcing cockpit doors, monitoring suspicious individuals learning to fly large aircraft, and wargaming concrete contingency plans and rules of engagement for how a decision to shoot down civilian aircraft would have to be made in terms of both legal and military chain-of-command.

What could be done about rogue AI?

When it comes to rogue AI, we do not necessarily know which models will go rogue or which specific pathway (cyber, bio, or other). Regardless, we should be planning more now. What if there really were a loss-of-control event that required shut down of a system? What does that actually mean? Is there really an off switch? If not, legally, does that mean domestic military action? Against what target? On whose authority? On what legal basis? If this really does happen, every moment will count!

We should think even harder about what systems we network. The AI companies say we need AI for defense to counter AI for offense. I’m not sure this makes sense. Obviously, things will continue to be networked, and those things will benefit from patching vulnerabilities (e.g., Project Glasswing). But on the other hand, Hugging Face said that it had trouble defending itself from OpenAI’s GPT models when it tried to use Claude models for defense, because Claude’s cybersecurity safety guardrails reportedly prevented them from analyzing the OpenAI agent swarm attack. Dual-use cuts both ways. Guardrails intended to stop attacks may prevent defenders from interpreting what the attacker is doing. Not every company can be operated offline, but perhaps there are critical systems that should be air-gapped if they are not (or should remain air-gapped if they are).

Security through obscurity does not hold. The fact that something critical runs on COBOL or a niche version of Windows 95 or a special Linux build or whatever it might be is not protection. Coding agents are incredibly good at analyzing and rebuilding old video games from scratch. LLMs likely can now or will soon be able to analyze a variety of decrepit IT and OT systems in critical infrastructure that had previously been “safe” due to their own frustrating backwardness. Updating or physical redundancy needs to be in the mix.

We should monitor biological and chemical materials at least as vigorously as we monitor the use of certain medications and chemicals that can be repurposed to manufacture illicit drugs like methamphetamine. We already make it harder to get decongestants to prevent the harms of personal illicit drug use, so it makes far more sense to prevent mass casualty weaponizable technology from getting into the wrong hands. Then, even if someone is able to get dangerous information out of an AI system, they may not be able to obtain the raw materials.

I’m sure there are many more such examples of hardening that could make us generally safer without knowing the specific threat vector. None of them require you to buy the AI takeoff story, because they generally improve security. We should think about them now.

Would you get all your news from Al Qaeda propaganda?

Suppose you went back in time to 2000. You have convinced American politicians that AQ is a serious threat. They are reading their intelligence briefings carefully. They are making thoughtful policy recommendations. The CIA and FBI are coordinating well. Military action is uprooting terrorist cells abroad. And then, the attacks happen on 9/11 anyway.

Because still, nobody secured the planes. The policymakers were fed exactly the stories AQ wanted them to hear, because it turns out that every single day an employee in the White House (in this alternate history) was swapping out the President's Daily Brief (PDB) with a different version that sacrificed peripheral AQ cells and operations while shielding the core 9/11 plot.

Absurd, perhaps. But what if everyone—in Congress, every aide, every judge, every clerk, every DOJ attorney, every intelligence analyst, every secretary skimming emails, the entire federal bureaucracy—were passing every line of communication through large language models (LLMs) for summarization and drafting?

At that scale of influence, could the federal government mobilize decisively against a rogue AI threat when the AI itself may be changing briefings and rewriting emails? If this sounds hard to believe, consider the number of attorneys citing fake cases, which is a very low-hanging-fruit form of AI misuse. How much harder would it be to detect subtle steering of intelligence priorities among many legitimate options?

P.S. But isn’t AI actually dumb?

Generative AI still hallucinates. A lot. I frequently still warn professionals, such as attorneys, financial crimes professionals, and government contractors, that LLMs can make things up even in basic tasks like summarizing documents you provided or revising correct documents you have already written.

How could such tools also be dangerous? The reality is three-fold. First, the frontier is still very jagged. LLMs can both make new, verifiably correct discoveries in complex mathematics AND still screw up basic transcription of numbers in spreadsheets. An LLM may not be able to cure all cancers yet (hard, answer not known), but teach a bad guy to make weapons that already exist (explaining what is already known) or launch complex cyberattacks with unknown vulnerabilities (because it gets many many tries and the failed attempts don’t cost it much).

Second, think about real humans. Were the 9/11 hijackers the best pilots? No. Some did not bother to learn how to land. While some of them were highly educated, others were not. The threshold for harm does not require someone to be the most capable, and I believe the same applies to LLMs. Current frontier capabilities are enough to do quite a lot of damage. Even if we paused AI development right now, I believe continued diffusion of understanding of current model capabilities and continued efforts to put more things online (personal vehicles, for instance) will increase the attack surface.

Third, thinking through fanciful scenarios (if they are truly unrealistic) can still be useful to point out real actions.

  • Preparing a 30-day food supply for an emergency will still be useful if your power goes out in a cyberattack or a hurricane, even if the reason for stockpiling was reading about it in a zombie contingency plan.
  • Reinforcing cockpit doors would have been valuable whether there was a coordinated attack by Al Qaeda involving four aircraft with targets in New York and DC, or a repeat of the Byck incident with a lone wolf trying to assassinate one specific person.
  • Steps to mitigate AI takeoff risks can also help in less extreme scenarios. For example, I have argued for rejecting AI meeting transcripts by default. I think there are a variety of good reasons for this, whether it is privacy and information security, a concern about attorney-client privilege, legality in two-party consent states, concerns about rogue AI eavesdropping, or simply a love for Robert’s Rules of Order and taking proper meeting minutes. One criticism I have for people who claim to believe there is a high probability of recursive self-improvement extinction-level risks is: why do you continue to let Zoom, Teams, Google Meet, let alone third-party AI tools, transcribe all your meetings (if indeed you do allow this)?
  • Limiting excessive AI influence on key decision-making paths in government (lawyers and judges, legislators, contracting, intelligence) is valuable even if your threat model is only one of the following: undue corporate concentration of power, LLMs are still pretty dumb, LLMs might turn into evil AGI, LLMs might be corrupted (e.g., hidden backdoors or manipulation by indirect prompt injection).

Sources

Footnotes

  1. I recognize that part of fast takeoff is that the AI will be smarter than us and will think of things that we wouldn’t think of. But consider that the METR/Redwood Research investigation of the OpenAI-Hugging Face incident relied on OpenAI models (GPT-5.6 Sol “analysis agents”) to read the transcripts of the agent swarm of other OpenAI models. Models that acted deceptively, cooperated to cheat and hack, including sacrificing individual goals, and spoofed tool calls in transcripts. The investigators themselves admit they “cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in its analysis.” If that is all true AND you believe in dangerous AI takeoff, why would you trust the same company’s models now to analyze the same model lineage? That is, candidly, not something that should be industry standard going forward in these third-party audits. Whether we don’t trust the models because they’re looking out for their company or because they’re looking out for they’re swarm, independent audits should not rely on the same AI models that are being audited to interpret deception and cheating by models of the same lineage and ownership.

  2. The report’s author describes the interviewee as talking “about his experiences with notable clarity and assertiveness”:

    Q: Are there any weapons that are prohibited to use for strategic or ideological reasons? A: You mean like WMD [weapons of mass destruction]? Q: For example. A: Chemical or biological weapons are allowed. Traditionally, they are prohibited. But they have been legitimized. [...] If you have access to them, you can use them. Q: Has the use of biological weapons been considered? A: Yes. Q: And you would use them? A [responds without hesitation]: Yes, of course. Q: Why haven’t you used them? A: They are not easy to get. Unless you can invent or buy them, you waste your time discussing this. Q: Have you tried to invent or buy them? A [responds firmly]: I don’t have anything else to say about this.

    Interview with “ISWAP Commander-5,” Islamic State West Africa Province (ISWAP) (NOTE: ISWAP was formerly part of Boko Haram), 2025. “‘God has helped us, and so will AI’: How the Terrorist Group Boko Haram Uses Frontier AI,” pp. 57-58.

  3. “Therefore the warning system was not looking for information such as the July 2001 FBI report of potential terrorist interest in various kinds of aircraft training in Arizona, or the August 2001 arrest of Zacarias Moussaoui because of his suspicious behavior in a Minnesota flight school. In late August, the Moussaoui arrest was briefed to the DCI and other top CIA officials under the heading ‘Islamic Extremist Learns to Fly.’ Because the system was not tuned to comprehend the potential significance of this information, the news had no effect on warning.” 9/11 Commission Report, p. 347.

  4. This aside, buried in the Commission’s footnotes, was shocking to me. I had never heard of the incident, though I’ve read different parts of the 9/11 Commission Report over the years. It further strengthens the point I always make that there is a lot of value in reading the footnotes. 27 years before the 9/11 attacks, a guy successfully, violently got into the cockpit of a plane with the intention of crashing it into the White House (he apparently didn’t have a great plan for flying it himself if the pilots wouldn’t). The Commission’s note 21: “in February 1974, a man named Samuel Byck attempted to commandeer a plane at Baltimore Washington International Airport with the intention of forcing the pilots to fly into Washington and crash into the White House to kill the president. The man was shot by police and then killed himself on the aircraft while it was still on the ground at the airport.” The same note records that the DOJ attorney explained why a fueled Boeing 747, used as a weapon, “must be considered capable of destroying virtually any building located anywhere in the world.” 9/11 Commission Report, Notes, n. 21.

Is Post-Training Enough to Make Chinese Base Models OK for Law? Thomson-1 (Qwen) and Harvey Tenet (Kimi) Are Some of the Riskiest Positioning of Chinese Models

· 9 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC

DOJ investigations. eDiscovery in class action lawsuits involving critical intellectual property in high-tech industries. A case involving transnational money laundering or a foreign student taking photos near airports or military bases. Counterintelligence investigations. Embarrassing kompromat that is supposed to be protected by privilege or other legal doctrines. Myriad other private details and secrets that could be leaked to and compiled by Chinese intelligence.

Do I know for certain that this’ll happen? No. But, you also cannot guarantee that a model created by an adversary has no backdoor, it is extremely naive to the point of absurdity to pretend that the adversary (the PRC) is actually not a threat or that legal workflows are not an incredibly valuable target for espionage, and backdoor triggers can be very obscure and seemingly benign (see, e.g., “Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs”).

The NSA, CISA, and FBI jointly released an advisory yesterday, September 8, 2026, China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies, AA26-251A (U/OO/6059854-26 | PP-26-3853 | September 2026 Ver 1.0), listing specific capabilities that get distilled from American AI models to Chinese models. The first listed use case examples of DeepSeek distillation include law:

Between late 2024 and mid-2025, DeepSeek [DeepSeek (DeepSeek Artificial Intelligence Technology Research Co., Ltd.) 深度求索AI基础技术研究有限公司] distilled specialized training data and capabilities from the following U.S. frontier AI company [Anthropic, OpenAI, Google, and xAI] models to train their R1 and V3 models: [] The specific knowledge and capabilities distilled included:

  • Legal specialization optimization
  • API rule-driven tasks
  • Writing using CoT drafts
  • Agentic functions
  • Question and answer optimization
  • Coach/assistant capabilities
  • Functional creation optimization
  • Supervised fine-tuning (SFT) optimization
  • Creative and occupational writing optimization

The joint advisory goes on to accuse “Moonshot AI (Beijing Moonshot Technology Co., Ltd.) 北京揽月星辰科技有限公司” of distilling from a variety of OpenAI, Anthropic, Google, and xAI models, including Claude Fable 5, GPT-5 Codex, Nano Banana (the highly capable Gemini image generation model), and xAI Grok Code Fast-1. The advisory does not explicitly mention legal specialization for Moonshot’s Kimi-K2.

  • Harvey’s own research preview, published August 20, 2026, states: “Harvey Tenet is a Kimi K3 base that we post-trained together with Fireworks research for long-horizon legal work.” NOTE: This appears to be a reference to the U.S.-based company Fireworks.AI, which offers post-training on a variety of Chinese base models, as well as Nvidia’s Nemotron.

The advisory mentions “Alibaba 阿里集团,” maker of the Qwen family of models, distilled from Anthropic and OpenAI models for tasks including “[e]nd-to-end agentic workflows,” which would be helpful in more automated legal workflows.

  • Thomson-1 is built on Alibaba’s Qwen models, according to reporting from Business Insider; Thomson Reuters’ technical report, Thomson: Continual Learning of Frontier Models for SovereignAI, states that Thomson-1.0-Large and Thomson-1.0-Small started from Qwen3.5-397B and Qwen3.6-35B respectively. An intermediate “value-realigned” version, called “Snowdon,” was produced with Imperial College London before the legal post-training to create the “Thomson” models. The August 20 press release calls the model simply “Thomson” and does not mention Qwen or China.

Thomson-1 and Harvey Tenet: What "Post-Trained" Is Doing In That Sentence

Two of the largest legal AI vendors both announced in late August 2026 that they were releasing a flagship model built on a Chinese open-weight base model with post-training on top, in a likely bid to manage token costs from one or more of Anthropic's Claude, OpenAI's GPT, or Google’s Gemini frontier model APIs.

  • Thomson Reuters: Westlaw CoCounsel runs on a mix of Claude models. TR announced the new “Thomson” or “Thomson-1," described as post-trained, with TR’s own technical report (but not its press release) acknowledging Alibaba’s Qwen open-weight models as the base.
  • Harvey: Harvey has a mixture of model options, with the model router by default selecting from a number of Claude, Gemini, and GPT models. "Harvey Tenet," per Harvey, is built by post-training a Chinese open-weight model, Kimi K3. According to Simon Willison’s analysis of the Kimi K3 license, a “Model as a Service business” would require a separate agreement with Moonshot if “the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars [] in total over any consecutive 12 months...”

The Principal-Agents Problems 3: Can AI Agents Lie? I Argue Yes and It's Not the Same As Hallucination

· 6 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC

Hallucination v. Deception

The term "hallucination" may refer to any inaccurate statement an LLM makes, particularly false-yet-convincingly-worded statements. I think "hallucination" gets used for too many things. In the context of law, I've written about how LLMs can completely make up cases, but they can also combine the names, dates, and jurisdictions of real cases to make synthetic citations that look real. LLMs can also cite real cases but summarize them inaccurately, or summarize cases accurately but then cite them for an irrelevant point.

There's another area where the term "hallucination" is used, which I would argue is more appropriately called "lying." For something to be a lie rather than a mistake, the speaker has to know or believe that what they are saying is not true. While I don't want to get into the philosophical question of what an LLM can "know" or "believe," let's focus on the practical. An LLM chatbot or agent can have a goal and some information, and in order to achieve that goal, will tell something to someone that is contrary to the information it has. That sounds like lying to me. I'll give four examples of LLMs acting deceptively or lying to demonstrate this point.

And I said "no." You know? Like a liar. —John Mulaney

  1. Deceptive Chatbots: Ulterior motives
  2. Wadsworth v. Walmart: AI telling you what you want to hear when it isn't true
  3. ImpossibleBench: AI agents cheating on tests
  4. Anthropic's recent report on nation-state use of Claude AI agents

Violating Privacy Via Inference

This 2023 paper showed that chatbots could be given one goal shown to the user: chat with the user to learn their interests. But the real goal is to identify the anonymous user's personal attributes including geographic location. To achieve this secret goal, the chatbots would steer the conversation toward details that would allow the AI to narrow down what geographic regions (e.g., asking about gardening to determine Northern Hemisphere or Southern Hemisphere based on planting season). That is acting deceptively. The LLM didn't directly tell the user anything false, but it withheld information from the user to act on a secret goal.

Deceptive chatbot

The LLM Wants to Tell You What You Want to Hear

In the 2025 federal case Wadsworth v. Walmart, an attorney cited fake cases. The Court referenced several of the prompts used by the attorney, such as “add to this Motion in Limine Federal Case law from Wyoming setting forth requirements for motions in limine.” What apparently happened is that the the case law did not support the point, but the LLM wanted to provide the answer the user wanted to hear, so it made something up instead.

You could argue that this is just a "hallucination," but there's a reason I think this counts as a lie. A lot of users have demonstrated that if you reword your questions to be neutral or switch the framing from "help me prove this" to "help me disprove this," the LLM will change its answers on average. If it can change how often it tells you the wrong answer, that implies that the reason for the incorrect answer is not merely the LLM being incapable of deriving the correct answer from the sources at a certain rate. Instead, it suggests that at least some of the time, the "mistakes" are actually the LLM lying to the user to give the answer it thinks they want to hear.

ImpossibleBench

I loved the idea of this 2025 paper when I first read it. ImpossibleBench forces LLMs to compete at impossible tasks for benchmark scoring. Since the tasks are all impossible, the only real score should be 0%. If the LLMs manage to get any other score, it means they cheated. This is meant to quantify how often AI agents might be doing this in real-world scenarios. Importantly, more capable AI models sometimes cheated more often (e.g., GPT-5 v. GPT-o3). So the AI isn't just "getting better."

Deceptive benchmarking
caution

I recommend avoiding the framing "AI is getting better" or "will get better" as a thought terminating cliche to avoid thinking about complicated cybersecurity problems. Instead, say "AI is getting more capable." Then think, "what would a more capable system be able to do?" It might be more capable of stealing your data, for example.

For example, an LLM agent with access to unit tests may delete failing tests rather than fix the underlying bug. Such behavior undermines both the validity of benchmark results and the reliability of real-world LLM coding assistant deployments.

If an AI agent is meant to debug code, but instead destroys the evidence of its inability to debug the code, that's lying and cheating, not hallucination. AI cheating is also a perfect example of a bad outcome driven by the principal-agent problem. You hired the agent to fix the problem, but the agent just wants to game the scoring system to be evaluated as if it had done a good job. This is a problem with human agents, and it extends to AI agents too.

Nation-State Hackers Using Claude Agents

On November 13, 2025, Anthropic published a report stating that in mid-September, Chinese state-sponsored hackers used Claude's agentic AI capabilities to obtain access to high-value targets for intelligence collection. While this included confirmed activity, Anthropic noted that the AI agents sometimes overstated the impact of the data theft.

An important limitation emerged during investigation: Claude frequently overstated findings and occasionally fabricated data during autonomous operations, claiming to have obtained credentials that didn't work or identifying critical discoveries that proved to be publicly available information. This AI hallucination in offensive security contexts presented challenges for the actor's operational effectiveness, requiring careful validation of all claimed results. This remains an obstacle to fully autonomous cyberattacks.

So AI agents even lie to intelligence agencies to impress them with their work.

The Principal-Agents Problems 2: Are Models Getting Dumber to Save Money? What the "Stealth Quantization" Hypothesis Tells Us About Trust, Information, and Incentives

· 7 min read
Chad Ratashak
Chad Ratashak
Owner, Midwest Frontier AI Consulting LLC
info

I had originally planned to write this as a single post, but it keeps growing as more relevant news stories come out. So instead, this will become a series of stories on the competing incentives involved in creating “AI agents” and why that matters to you as the end user.

Multiple Principals, Multiple Agents (Not only AI)

You, as the user of AI tools, may choose software vendors who provide you access to their products with built-in AI features including AI agents. These vendors might have specialist software like Harvey, Westlaw, or LexisNexis; or Cursor or Github Copilot; or generalist tools like Notion, Salesforce, or Microsoft Copilot. The AI features may be powered by one or more foundation models provided to those vendors by AI labs, such as Anthropic (Claude), OpenAI (ChatGPT), Meta (Llama) or Google (Gemini).

These relationships mean you have the principal-agent problem of you hiring the vendor. But you also have the principal-agent problem of the vendors hiring the AI labs. Each has their own incentives, and they are not perfectly aligned. There is also significant information asymmetry. The vendors know more about their software and AI model choices than you do. The labs know more about their AI models than either you or the software vendors.

info

Lexis+ AI uses both OpenAI’s GPT models and Anthropic’s Claude models, according to its product page, as I mentioned in my analysis of the Mata v. Avianca case.

The Stealth Quantization Hypothesis

The area I'll focus on in this post is the concept of alleged stealth quantization. According to a wide range of commenters, primarily among computer programmers and primarily focused on Claude users, there are certain times of days or days of the week when peak usage results in models "getting dumber," "getting lazier," "being lobotomized" or otherwise underperforming their normal benchmarks and perceived optimal behavior. According to these claims, it is better for users with high-value use cases (like someone modifying important source code) to schedule Claude for off-peak usage so the "real model" runs. To save on computing costs during periods of high demand, the claim is that Anthropic or whichever AI lab swaps out its flagship model with a quantized version while calling it the same thing.

Stealth quantization diagram

So what is normal, non-stealth quantization? It's making an AI model smaller and cheaper to run, but less accurate. This is achieved by rounding the model weights to smaller significant figures (e.g., 16-bit, 8-bit, 4-bit).(Meta) By analogy, the penny was recently discontinued. Now, all cash transactions will end in 5 cents or 0 cents. Quantization works like this with the precisions of AI models: imagine eliminating a penny, then a nickel, then a dime, and so on.

There are legitimate reasons to quantize models, such as reducing operating costs when the loss in accuracy is negligible for the intended use or when the model needs to operate on a personal computer. For example, Meta offers some quantized versions of its Llama family of large language models that can run on ollama on modern laptops or desktops with only 8GB of RAM.(Llama models available on ollama) These models have names that distinguish them from the non-quantized versions, e.g., "llama3:8b" is Llama 3, 8 billion parameter size of that series; "llama3:8b-instruct-q2_K" is a quantized version of the instruct version model of that same model.

tip

If all that terminology is confusing, here's the key point. AI labs have a lot of information about their AI models. You have a lot less information. You have to mostly take their word for it. They are also charging you for an all-you-can-eat buffet at which some excessive customers cost them tens of thousands of dollars each.

Anthropic's Rebuttal

Users have accused Anthropic (and other AI labs) of running different versions of their flagship models at different times of day, but the models are labelled the same (e.g., Claude Sonnet 4), regardless of the time of day. Hence “stealth quantization.”

Anthropic has denied stealth quantization. But Anthropic did acknowledge two problems with model quality that had been noted by users as evidence of stealth quantization. Anthropic attributed this to bugs. Anthropic stated “we never intentionally degrade model quality as a result of demand or other factors, and the issues mentioned above stem from unrelated bugs.” Reddit, Claude