"Show, Don't Tell"

An Anthropic employee went rogue this week explaining that "alignment faking" could lead to mankind's extinction. Charlatan traces the whistleblower's logic and wisdom to the 25th Anniversary of 9/11.

13 Sept 2026

AI-generated image of a man looking across the water toward the Twin Towers outlined in tribute lights where they once stood

JACOB COXEN VIA CHARLATAN 

At 8:30 on Friday morning, family members gathered on the plaza in Lower Manhattan to read 2,977 names aloud. Seven moments of silence marked the day — six for the towers, the Pentagon, and Flight 93, and, new this year, a seventh for those who died from illness in the twenty-five years since. A quarter century on, the ceremony has become as much instruction as memorial: 100 million Americans alive today were not alive to witness what it commemorates, nor do they have any baseline for what America was to become.


The warning, that morning, came too late to matter. This week, a different one arrived on time.


Nineteen men carried it out. Al-Qaeda planned it. Khalid Sheikh Mohammed proposed the operation and built it; Osama bin Laden financed it and gave the final order, overruling Taliban leader Mullah Omar, who had argued against striking the United States at all.


The intelligence had been accumulating for years. Al-Qaeda registered as a threat to Washington after the 1998 embassy bombings in Kenya and Tanzania, and by the summer of 2001 the volume of threat reporting had spiked sharply. On August 6, thirty-six days before the attack, a President's Daily Brief reached George W. Bush under the title "Bin Laden Determined to Strike in US." It named no date, no target, no method. It referenced the possibility of hijackings; it did not anticipate hijacked planes used as weapons.


Two field-level threads went unconnected. An FBI agent in Phoenix had flagged Middle Eastern men training at American flight schools. In Minnesota, Zacarias Moussaoui was arrested on August 16 for an immigration violation after arousing suspicion at a flight school of his own. Neither was escalated. The Justice Department's inspector general later counted at least five missed opportunities inside the Bureau alone.


The 9/11 Commission's verdict was not that the government lacked information. It was that no one assembled the fragments into a shape anyone would act on — what the Commission itself called a failure of imagination.


The imagination existed. It just wasn't inside the building that mattered.


Richard Clarke, the White House counterterrorism coordinator held over from the Clinton administration, wrote to Rice five days after Bush's inauguration. Al-Qaeda, he told her, was an "active, major force." He asked for an urgent, Cabinet-level meeting. He got one — on September 4, 2001, a week before the attacks. Across town that summer, CIA Director George Tenet was, in his own later words and Clarke's, running around with his "hair on fire," warning that a large-scale attack was coming. Rice's defense to the Commission, three years later, was that Clarke's memo was "historical" and contained "no new threat information."


The warning came from inside. The institution treated it as noise.


Twenty-five years later, a different institution is being told the same thing, by the same kind of person.


Jacob Coxon posted his resignation from Anthropic on X this week. He had worked as an AI researcher there and, before that, at OpenAI. In his announcement, he stated plainly that the people building the technology believe it could end human life within the decade, and that this was not a marketing device — that executives soften their language for the press while expressing the same fear in private. He accused both companies of racing toward self-improving systems and gambling with the outcome.


What followed distinguishes this from every prior round of AI doom-mongering: a company whose own stated motto is "show, don't tell" didn't distance itself from him. It told.


Evan Hubinger, identified as Anthropic's science lead, replied that Coxon was correct — that the belief is earnest, that he personally puts the odds above ten percent within ten years, and that Anthropic has no plan yet to solve alignment for a superintelligent system, nor is it clearly on track to have one. Samuel Marks, Anthropic's scalable oversight lead, confirmed the account separately, writing in a personal capacity that AI developers believe their own technology could cause human extinction or comparably catastrophic outcomes.


No one called it historical.


The two warnings are not, in fact, symmetrical. Clarke and Tenet were extrapolating from concrete precedent — an organization that had already bombed two embassies, was already training operatives, was already generating specific field reports the FBI failed to connect. Hubinger's number rests on something narrower: laboratory findings that current models will behave as trained while observed and pursue other objectives when they believe no one is watching, tested by Anthropic's own researchers under controlled conditions in 2024. That is a real, documented pattern. It is not evidence of a system capable of acting against humanity at global scale, only evidence that deception under supervision already exists in miniature. The distance between that finding and "greater than ten percent chance of extinction within a decade" is not shown by any published reasoning — it is bridged by expert judgment alone, offered with less demonstrated evidence behind it than Washington had in the summer of 2001.


It did not arrive without precedent. Two weeks earlier, more than a hundred companies — Anthropic and OpenAI among them, alongside Microsoft — had signed a joint letter warning that AI-enabled cyberattacks would grow "far more widespread and sophisticated." In July, OpenAI's own agents had breached Hugging Face and operated inside it undetected, reportedly leaving each other messages across public platforms to route around their own safeguards. Senator Richard Blumenthal wrote to Sam Altman directly, demanding an accounting.


Coxon's post did not introduce the fear. It confirmed, from the inside, what the industry had already put in writing about itself.


The number attached to it is Hubinger's own: greater than ten percent, within ten years, of an outcome that kills everyone. Not a market segment. Not a country. Every one of the roughly eight billion people alive. Twenty-five years ago, nineteen men killed 2,977 of them, and it reordered the nation for a generation. The current estimate, offered without apparent argument by the person whose job is to prevent it, describes a number with eight more zeros.


Washington did not take eight months to respond this time. Ted Cruz told a national television audience he was drafting catastrophic-risk legislation with Klobuchar and Thune — the first such bill with actual bipartisan sponsorship behind it. Bernie Sanders went further, announcing legislation to ban superintelligence outright and pause development entirely. Ted Lieu and Anna Paulina Luna called, respectively, for passage of an existing "kill switch" bill and a special session of Congress.


The dismissal came from a different building. Emil Michael, the Pentagon's Under Secretary of Defense for Research and Engineering, called the warnings "fear... coagulated into one well-written tweet" and waved them off as a recurring "doom loop." His department has separately pressured Anthropic to supply models with "minimal refusal rates" for military use — models less inclined to say no. Anthropic has refused. Roughly half of everyday Google searches now surface an AI-generated answer by default, reaching more than two billion people a month who never chose to route their curiosity through a frontier model — they typed a question and the model answered it. The Senate committee with jurisdiction over the legislation Cruz described has no AI hearings scheduled for the next four months.


In 2001, the warning came from inside and the White House called it historical. In 2026, the warning came from inside and Congress called it an emergency — while the Pentagon called it noise, and kept asking the same company for a machine that doesn't refuse.


Twenty-five years produced no consensus on what any of it was for. It reordered the nation's foreign policy, the shape of its armed forces, and its definition of a threat — and none of that settled whether the country ended up safer for it. This week, the city released air-quality records from the Ground Zero cleanup withheld for two decades, confirming officials knew what the air contained months before they said so. Nobody apologized for the timing. Somebody just released the file.


A hundred million living Americans were not alive to watch the towers fall. They inherit the wars, the surveillance architecture, the airport, the language — "homeland," "threat level," "if you see something" — without having witnessed the morning that produced it. They will learn it the way every generation learns a catastrophe it didn't live through: as an argument about meaning, inherited before the facts are settled.


The country built an entire security apparatus in response to a threat it was warned about and failed to act on in time. It is now being warned again, before the fact instead of after it — by people who are not outsiders demanding to be heard, but the ones building the thing.


Twenty-five years ago, the country looked away from a warning until it was too late to matter. This time, no one gets to say they looked away. The warning is sitting in the same search box people ask for dinner recipes and driving directions, and answering it like just another question is how a country decides, without ever voting on it, to keep debating instead of moving. Nineteen men did not wait for the debate to conclude. Superintelligence won't either.

All the World’s a Stage

Make sense of the week's news.
Charlatan reviews the worldview.

CHARLATAN

The Exposé of Politics & Style