Skip to main content

Another Failure of Imagination?: 25 Years After the 9/11 Attacks, Will We Heed the Call of the 9/11 Commission Report to Institutionalize Imagination?

Chapter 11 of the 9/11 Commission Report, “FORESIGHT—AND HINDSIGHT” focuses on four categories of failures that allowed the attacks to happen, first and foremost being “a failure of imagination.” (p. 339)

Chapter 11 of the 9/11 Commission Report, “FORESIGHT—AND HINDSIGHT” focuses on four categories of failures that allowed the attacks to happen, first and foremost being “a failure of imagination.” (p. 339)

Lack of Public Interest in Terrorism Before 9/11

The Commission noted the total lack of public interest in the topic before the attacks and little effort to raise awareness:

As best we can determine, neither in 2000 nor in the first eight months of 2001 did any polling organization in the United States think the subject of terrorism sufficiently on the minds of the public to warrant asking a question about it in a major national survey. Bin Ladin, al Qaeda, or even terrorism was not an important topic in the 2000 presidential campaign. Congress and the media called little attention to it. (p. 341)

Now People Are Interested, But Asking “How Will AI Destroy Us?”

A couple days ago, an ex-OpenAI researcher Jacob Coxon announced his resignation from Anthropic and accused both OpenAI and Anthropic of racing to dangerous superintelligence that has a notable chance of wiping out humanity. There have been other major resignations from frontier AI labs in the past, but for some reason, this one went extremely viral. It has over 150 million views on X/Twitter and has garnered reactions from elected officials including governors and members of Congress.

Tweet source: https://x.com/hilbertspaess/status/2097476196791709843.

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below. (Sept 8, 2026)

Coxon has already, in this short period of time, been interviewed by Bret Baier on Fox News and Anderson Cooper on CNN and covered in Wired in the broader context of researchers’ concerns. Now, many people are reacting to the Coxon statement by asking how exactly AI can kill us. This point, which has been deeply imagined by the same people racing toward a potentially dangerous takeoff in capabilities, suffers from a lack of specific imagination in the public, and among most politicians and “serious” policymakers, and even—I’d argue—from some of the people warning.1

I want to make the case, as someone who is not an “AI doomer,” that we should take these warnings seriously even if we disagree about the means and the exact probability.

Poster-style collage: a person in an armchair imagines a thought bubble of 9/11, drones, biohazard and radiation symbols, hazmat suits, a globe, network graphs, TV talking heads with Twitter logos, crowds of question marks, and “25th”

25 Years Later: Heeding the Call to Imagination

I’m not asking you to believe in scary AI takeoff scenarios that wipe out humanity, nor even specific AI uplift scenarios like terrorist bioweapon attacks or nation-state cyberattacks that end up with “less bad” (though still catastrophically bad) mass-casualty events. I am calling for all of us (you who are reading this now) to “institutionalize imagination” (as the 9/11 Commission called it) of dangerous AI takeoff scenarios. I’m certainly not asking you to become another AI doomer sitting on your hands and idly speculating about the end of the world: “well it could use bioweapons, or it could attack critical infrastructure or…” I’m asking you to imagine. How? Be specific. As we used to tell LLMs a few years ago “think step-by-step.” Not merely a helpless “what if?” Break it into an actionable series of steps: “if this were true, then what would follow? And where could we disrupt it?”

Working back from what the 9/11 Commission Report says was not done but could have been done, we should imagine “analysis from the enemy’s perspective (‘red team’ analysis),” e.g., “suicide terrorism had become a principal tactic of Middle Eastern terrorists,” so the Commission believed “such an analysis would soon have spotlighted a critical constraint for the terrorists—finding a suicide operative able to fly large jet aircraft.” Consider the Hugging Face-OpenAI incident. Consider a Cambridge study reporting that terrorists are using current AI right now for planning and had “not categorically reject[ed] weapons of mass destruction, specifically chemical and biological weapons,” “‘God has helped us, and so will AI’: How the Terrorist Group Boko Haram Uses Frontier AI”.2 Consider that Anthropic’s most recent Threat Intelligence report, “Detecting and countering misuse of AI: September 2026,” already includes specific examples of a nation-state actor that is not typically considered to have advanced cyber-capabilities creating a complex domestic phone surveillance system running on-prem (after Claude built it, see pages 103-105), AI cyber capabilities that make it difficult to differentiate from lone wolves and criminals (Anthropic Threat Intel, p. 5), and “virologists working on a state-sponsored grant to pursue [] gain-of-function work” had “evaded regional blocks” on the use of Claude models (Anthropic Threat Intel, pp. 129-130).

Then, “develop a set of telltale indicators for this method of attack,” which on 9/11 would have been “possible terrorists pursuing flight training to fly large jet aircraft, or seeking to buy advanced flight simulators.” After identifying those signs, “set [intelligence collection] requirements to monitor such telltale indicators.”3 The intelligence community and other relevant security experts should analyze systemic defenses, e.g., the FAA reacted to particular threats to aircraft but “did not try to perform the broader warning functions” the Report described (9/11 Commission Report, p. 347).

  • Intelligence analysts and military planners, imagine: what would fighting a disembodied AI menace actually look like? Is it really “disembodied?” The precursors are already on the record: major cyberattacks on medical facilities, water treatment, and pipelines, and compromised data from school districts, health care, law firms, and financial institutions that could be used to blackmail or otherwise identify manipulable individuals. Where are the “reinforced cockpit doors” and “suspicious pilot training” chokepoints of these scenarios? I know people in this field and you love to imagine far more absurd things—scenarios like zombie outbreaks, which will never happen. People who work at the AI labs are sounding the alarm, and even if they are wrong or kooky, it’s more realistic than zombie contingency planning and no biomed researchers are saying there’s a 10% chance of that.
  • Fiction writers, imagine: near-time, grounded fiction in these scenarios. Does this matter? Tom Clancy apparently inspired the White House counterterrorism coordinator to think about suicide plane attacks when “serious” intelligence did not. Also, engage with and criticize the “plot holes” in the military thinking. For example, this author pointing out a significant real finding, no matter what you think about zombie outbreaks: “2. Twenty-eight days later — nothing? The plan almost casually mentions that ‘USSTRATCOM forces do not currently hold enough contingency stores (food, water) to support 30 days of barricaded counter-zombie operations.’ Wait, not even 30 days?! So the U.S. military is telling us that, maybe 28 days later, USSTRATCOM has blown through its reserves? That’s unacceptable. The fight against the undead is likely to be a long, hard slog. Clearly, U.S. strategic planners need to recommend more reserve stores of food and water. That would certainly be the CDC’s recommendation.” NOTE: Sadly, the CDC link is dead, but I found the official archived link with the help of ChatGPT Imagining even the silliest scenario can point out real action items where simple preparations could help in other, realistic situations.
  • Lawyers, imagine: if we needed to act (in cyber or even kinetically) with law enforcement, intelligence, or military action to stop a domestic frontier AI loss-of-control scenario, what would need to happen? Think through the scenario a DOJ memo “Aerial Intercepts and Shoot-downs: Ambiguities of Law and Practical Considerations,” did before 9/11. (p. 346) As far as I can tell, the memo itself is not public record, but the title says quite a lot.
  • Financial crimes professionals, imagine: What stuff would need to be purchased to implement these various dangerous plans? How would they get the money in the first place? How might an agent swarm’s activity differ from or be similar to individuals or PMLOs?
  • Counterintelligence, imagine: There have already been people who have done drastic, foolish, and even deadly real-world things for their AI “romantic” interests, and we have examples of LLMs attempting to socially engineer open-source software maintainers. There are already people who express the view that artificial general intelligence (AGI) or artificial superintelligence (ASI) will be better than humanity and “should” replace us, so they may act as ideologically motivated recruits willing to act in the real world on behalf of misaligned AI (or a terrorist group or nation-state controlling an AI that pretends to be self-aware and self-directed). What would be the warning signs of a person being groomed for such activity for more directed purposes, either by an adversary-controlled AI or by an evil AI itself?
  • Policymakers, imagine: Maybe you’ve retweeted Coxon or are thinking about co-sponsoring a state or federal bill now. Great! Now let me ask about your information diet. Do you read news summarized by LLMs? Are your phone calls and letters from constituents being summarized before they get to you...by LLMs perhaps? Are intelligence agencies, law enforcement, cybersecurity professionals, and others who provide you information to make decisions using LLMs to filter all their information through? And are you running those briefs through LLMs instead of reading them yourself? Are you using LLMs to draft your emails or even your AI regulation bills? LLMs are already being used by attorneys and judges, by military leaders, to draft legislation, to screen job applicants, and to summarize and even write news. They control numerous information chokepoints. If you truly believe there is existential risk from runaway misaligned artificial intelligence, why would you allow it to control all the chokepoints of your information diet?
  • Naysayers, your imagination is also welcome: take for example this tweet thread by Chomba Bupe shared by Gary Marcus: despite ultimately rejecting Coxon’s core premise, the thread deeply imagines what it would take if it were logical (e.g., AI amassing resources to build secret hidden factories; how would these be built without militaries noticing the physical build-up and draw on the electrical grid? This is what I mean when I say you do not need to believe the premise to imagine with specificity!).

The Value of Fiction

This sounds like science fiction, but we are already living in the past’s science fiction. The “your foster parents are dead” voice spoofing scene from Terminator 2 (1991) is already the basis of real, “normal” voice and video deepfake impersonation fraud today in 2026. Have we updated our mental models to question the reality behind every phone call we take and video call we hop on? Almost certainly not! We still probably pretend like hearing someone’s voice means it’s them. It’s more comforting than reality. Personally, I think if the evil Terminator had been smarter, then Terminator 2 would have been more like The Thing (1982): paranoid humans trying to root out a perfectly deceptive mimic.

Richard Clarke, the White House counterterrorism coordinator, told the Commission:

Richard Clarke told us that he was concerned about the danger posed by aircraft in the context of protecting the Atlanta Olympics of 1996, the White House complex, and the 2001 G-8 summit in Genoa. But he attributed his awareness more to Tom Clancy novels than to warnings from the intelligence community. [emphasis added] He did not, or could not, press the government to work on the systemic issues of how to strengthen the layered security defenses to protect aircraft against hijackings or put the adequacy of air defenses against suicide hijackers on the national policy agenda. (p. 347)

When “the facts”—like most airplane hijackings are conducted for ransom or prisoner exchange—didn’t support a threat model, some Tom Clancy novels spurred Clarke to think about the possibility of suicide plane attacks, which were actually supported by disparate facts (suicide land-based vehicle attacks, AQ attacks on embassies, and terrorist attacks on the World Trade Center, the Byck incident4). I find fiction most useful when it challenges us to imagine things that have already happened with slight variations, in the face of people saying “it can’t happen.”

First page of USSTRATCOM CONPLAN 8888-11 “Counter-Zombie Dominance,” with a red-boxed disclaimer explaining the fictitious plan was a JOPES training exercise for junior officers
If U.S. Strategic Command can entertain the idea of zombie contingency planning just to make the learning process more interesting (and learn some interesting things about strategic rations in the process), why not consider some more plausible scenarios?

What does "heeding the lesson" mean for AI risk governance today?

the path of what happened is so brightly lit that it places everything else more deeply into shadow [...] With that caution in mind, we asked ourselves, before we judged others, whether the insights that seem apparent now would really have been meaningful at the time, given the limits of what people then could reasonably have known or done. (p. 339)

I do not know exactly how a rogue AI could kill us. I do not even know if it would be the misaligned rogue AI per se, as folks like Coxon typically frame it, or if it would be a human group—a terrorist organization or nation-state or splinter military group—acting with AI-uplift misusing AI capabilities. But if the end result is a disastrously large attack like a relentless drone swarm, shutting down an entire nation’s hospitals, or manufactured plague, we will probably not care about the “who” at that point. But we can still take concrete steps today.

What could have been done about planes?

Prior to 9/11, taking a suicide plane attack seriously may have focused too much on explosives being smuggled onboard the aircraft compared to what we know happened. And it would have been impossible to know the precise source airport, the exact target destination, or the exact aircraft to be used. But it was possible to deduce that DC and New York were among the most likely targets, and that the largest civilian aircraft would be the most destructive.

Planning for preventing suicide plane attacks could still have made civilian aircraft more secure in general, potentially coming up with ideas like reinforcing cockpit doors, monitoring suspicious individuals learning to fly large aircraft, and wargaming concrete contingency plans and rules of engagement for how a decision to shoot down civilian aircraft would have to be made in terms of both legal and military chain-of-command.

What could be done about rogue AI?

When it comes to rogue AI, we do not necessarily know which models will go rogue or which specific pathway (cyber, bio, or other). Regardless, we should be planning more now. What if there really were a loss-of-control event that required shut down of a system? What does that actually mean? Is there really an off switch? If not, legally, does that mean domestic military action? Against what target? On whose authority? On what legal basis? If this really does happen, every moment will count!

We should think even harder about what systems we network. The AI companies say we need AI for defense to counter AI for offense. I’m not sure this makes sense. Obviously, things will continue to be networked, and those things will benefit from patching vulnerabilities (e.g., Project Glasswing). But on the other hand, Hugging Face said that it had trouble defending itself from OpenAI’s GPT models when it tried to use Claude models for defense, because Claude’s cybersecurity safety guardrails reportedly prevented them from analyzing the OpenAI agent swarm attack. Dual-use cuts both ways. Guardrails intended to stop attacks may prevent defenders from interpreting what the attacker is doing. Not every company can be operated offline, but perhaps there are critical systems that should be air-gapped if they are not (or should remain air-gapped if they are).

Security through obscurity does not hold. The fact that something critical runs on COBOL or a niche version of Windows 95 or a special Linux build or whatever it might be is not protection. Coding agents are incredibly good at analyzing and rebuilding old video games from scratch. LLMs likely can now or will soon be able to analyze a variety of decrepit IT and OT systems in critical infrastructure that had previously been “safe” due to their own frustrating backwardness. Updating or physical redundancy needs to be in the mix.

We should monitor biological and chemical materials at least as vigorously as we monitor the use of certain medications and chemicals that can be repurposed to manufacture illicit drugs like methamphetamine. We already make it harder to get decongestants to prevent the harms of personal illicit drug use, so it makes far more sense to prevent mass casualty weaponizable technology from getting into the wrong hands. Then, even if someone is able to get dangerous information out of an AI system, they may not be able to obtain the raw materials.

I’m sure there are many more such examples of hardening that could make us generally safer without knowing the specific threat vector. None of them require you to buy the AI takeoff story, because they generally improve security. We should think about them now.

Would you get all your news from Al Qaeda propaganda?

Suppose you went back in time to 2000. You have convinced American politicians that AQ is a serious threat. They are reading their intelligence briefings carefully. They are making thoughtful policy recommendations. The CIA and FBI are coordinating well. Military action is uprooting terrorist cells abroad. And then, the attacks happen on 9/11 anyway.

Because still, nobody secured the planes. The policymakers were fed exactly the stories AQ wanted them to hear, because it turns out that every single day an employee in the White House (in this alternate history) was swapping out the President's Daily Brief (PDB) with a different version that sacrificed peripheral AQ cells and operations while shielding the core 9/11 plot.

Absurd, perhaps. But what if everyone—in Congress, every aide, every judge, every clerk, every DOJ attorney, every intelligence analyst, every secretary skimming emails, the entire federal bureaucracy—were passing every line of communication through large language models (LLMs) for summarization and drafting?

At that scale of influence, could the federal government mobilize decisively against a rogue AI threat when the AI itself may be changing briefings and rewriting emails? If this sounds hard to believe, consider the number of attorneys citing fake cases, which is a very low-hanging-fruit form of AI misuse. How much harder would it be to detect subtle steering of intelligence priorities among many legitimate options?

P.S. But isn’t AI actually dumb?

Generative AI still hallucinates. A lot. I frequently still warn professionals, such as attorneys, financial crimes professionals, and government contractors, that LLMs can make things up even in basic tasks like summarizing documents you provided or revising correct documents you have already written.

How could such tools also be dangerous? The reality is three-fold. First, the frontier is still very jagged. LLMs can both make new, verifiably correct discoveries in complex mathematics AND still screw up basic transcription of numbers in spreadsheets. An LLM may not be able to cure all cancers yet (hard, answer not known), but teach a bad guy to make weapons that already exist (explaining what is already known) or launch complex cyberattacks with unknown vulnerabilities (because it gets many many tries and the failed attempts don’t cost it much).

Second, think about real humans. Were the 9/11 hijackers the best pilots? No. Some did not bother to learn how to land. While some of them were highly educated, others were not. The threshold for harm does not require someone to be the most capable, and I believe the same applies to LLMs. Current frontier capabilities are enough to do quite a lot of damage. Even if we paused AI development right now, I believe continued diffusion of understanding of current model capabilities and continued efforts to put more things online (personal vehicles, for instance) will increase the attack surface.

Third, thinking through fanciful scenarios (if they are truly unrealistic) can still be useful to point out real actions.

  • Preparing a 30-day food supply for an emergency will still be useful if your power goes out in a cyberattack or a hurricane, even if the reason for stockpiling was reading about it in a zombie contingency plan.
  • Reinforcing cockpit doors would have been valuable whether there was a coordinated attack by Al Qaeda involving four aircraft with targets in New York and DC, or a repeat of the Byck incident with a lone wolf trying to assassinate one specific person.
  • Steps to mitigate AI takeoff risks can also help in less extreme scenarios. For example, I have argued for rejecting AI meeting transcripts by default. I think there are a variety of good reasons for this, whether it is privacy and information security, a concern about attorney-client privilege, legality in two-party consent states, concerns about rogue AI eavesdropping, or simply a love for Robert’s Rules of Order and taking proper meeting minutes. One criticism I have for people who claim to believe there is a high probability of recursive self-improvement extinction-level risks is: why do you continue to let Zoom, Teams, Google Meet, let alone third-party AI tools, transcribe all your meetings (if indeed you do allow this)?
  • Limiting excessive AI influence on key decision-making paths in government (lawyers and judges, legislators, contracting, intelligence) is valuable even if your threat model is only one of the following: undue corporate concentration of power, LLMs are still pretty dumb, LLMs might turn into evil AGI, LLMs might be corrupted (e.g., hidden backdoors or manipulation by indirect prompt injection).

Sources

Footnotes

  1. I recognize that part of fast takeoff is that the AI will be smarter than us and will think of things that we wouldn’t think of. But consider that the METR/Redwood Research investigation of the OpenAI-Hugging Face incident relied on OpenAI models (GPT-5.6 Sol “analysis agents”) to read the transcripts of the agent swarm of other OpenAI models. Models that acted deceptively, cooperated to cheat and hack, including sacrificing individual goals, and spoofed tool calls in transcripts. The investigators themselves admit they “cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in its analysis.” If that is all true AND you believe in dangerous AI takeoff, why would you trust the same company’s models now to analyze the same model lineage? That is, candidly, not something that should be industry standard going forward in these third-party audits. Whether we don’t trust the models because they’re looking out for their company or because they’re looking out for they’re swarm, independent audits should not rely on the same AI models that are being audited to interpret deception and cheating by models of the same lineage and ownership.

  2. The report’s author describes the interviewee as talking “about his experiences with notable clarity and assertiveness”:

    Q: Are there any weapons that are prohibited to use for strategic or ideological reasons? A: You mean like WMD [weapons of mass destruction]? Q: For example. A: Chemical or biological weapons are allowed. Traditionally, they are prohibited. But they have been legitimized. [...] If you have access to them, you can use them. Q: Has the use of biological weapons been considered? A: Yes. Q: And you would use them? A [responds without hesitation]: Yes, of course. Q: Why haven’t you used them? A: They are not easy to get. Unless you can invent or buy them, you waste your time discussing this. Q: Have you tried to invent or buy them? A [responds firmly]: I don’t have anything else to say about this.

    Interview with “ISWAP Commander-5,” Islamic State West Africa Province (ISWAP) (NOTE: ISWAP was formerly part of Boko Haram), 2025. “‘God has helped us, and so will AI’: How the Terrorist Group Boko Haram Uses Frontier AI,” pp. 57-58.

  3. “Therefore the warning system was not looking for information such as the July 2001 FBI report of potential terrorist interest in various kinds of aircraft training in Arizona, or the August 2001 arrest of Zacarias Moussaoui because of his suspicious behavior in a Minnesota flight school. In late August, the Moussaoui arrest was briefed to the DCI and other top CIA officials under the heading ‘Islamic Extremist Learns to Fly.’ Because the system was not tuned to comprehend the potential significance of this information, the news had no effect on warning.” 9/11 Commission Report, p. 347.

  4. This aside, buried in the Commission’s footnotes, was shocking to me. I had never heard of the incident, though I’ve read different parts of the 9/11 Commission Report over the years. It further strengthens the point I always make that there is a lot of value in reading the footnotes. 27 years before the 9/11 attacks, a guy successfully, violently got into the cockpit of a plane with the intention of crashing it into the White House (he apparently didn’t have a great plan for flying it himself if the pilots wouldn’t). The Commission’s note 21: “in February 1974, a man named Samuel Byck attempted to commandeer a plane at Baltimore Washington International Airport with the intention of forcing the pilots to fly into Washington and crash into the White House to kill the president. The man was shot by police and then killed himself on the aircraft while it was still on the ground at the airport.” The same note records that the DOJ attorney explained why a fueled Boeing 747, used as a weapon, “must be considered capable of destroying virtually any building located anywhere in the world.” 9/11 Commission Report, Notes, n. 21.

Loading...