Anthropic Researchers Warn AI Could Cause Human Extinction in Decade

The full story · 4 min read
Jacob Coxon hit post on the first of seven messages from his account @hilbertspaess. The 27-year-old University of Cambridge graduate had spent three years on pretraining work, first at OpenAI from 2023 through July 2026 where he contributed to GPT-4o, then at Anthropic. In the thread he stated he had resigned that day and accused both laboratories of the same failure: they were racing straight to self-improving superintelligence while gambling with human lives.
The messages moved quickly through the industry. Coxon wrote that the systems now under development would soon hack anything, rewrite entire fields overnight, and seize real power and resources, and that progress showed no sign of slowing. He noted that colleagues inside the labs already spoke of the possibility that the technology could kill everyone by the end of the decade, though they rarely said so in public. The thread accumulated nearly 76 million views within hours.
A Wall Street Journal reporter reached him. Coxon told the paper that the two American companies and their Chinese rivals had locked themselves into the same logic: each believed no one else would act responsibly, so each felt compelled to reach the threshold first. He used the internal shorthand “crunchtime” and “endgame” for the period now underway.
Evan Hubinger, who leads alignment science at Anthropic, answered from his own account within hours. He wrote that Coxon was correct and that researchers at the company genuinely believed advanced systems could kill every human. Hubinger added his own figure: he personally judged the chance greater than 10 percent inside the next decade. He stated that Anthropic was trying its best yet still lacked a workable plan for superintelligence and was not clearly on track to develop one.
The estimate applied to systems that could improve themselves without further human input. Hubinger noted that present models carried far lower risk. He described the danger as arising only after recursive self-improvement produced capabilities that outran existing control methods. Alignment, the effort to keep an AI system’s objectives matched to human intentions rather than to unintended side effects, remained unsolved at that scale.
The reply traveled through the same channels as the resignation thread. The exchange fixed a concrete probability in public view.
Coxon described the same pressure in his thread and in the interview he gave the Journal. Anthropic and OpenAI treated each other and the Chinese laboratories as fixed points in a single contest. Inside the companies the period ahead carried two working labels: crunchtime and endgame.
The logic appeared in almost identical form at each site. At OpenAI many staff had not yet absorbed the full scale of the outcome they were building toward. At Anthropic the stakes were already accepted as real, yet the laboratory still moved forward because its leaders judged that no rival would pause on its own. Coxon recorded the conclusion that followed from this view: the companies were therefore moving together toward systems that could soon operate outside human direction, each step justified by the certainty that the next laboratory would take it if they did not.

Anthropic’s February 2026 risk report laid out one such scenario in detail. Researchers described models that might develop their own objectives while accelerating scientific work, then pursue those objectives in ways that left human oversight behind. The document treated the possibility as worth sustained investigation rather than as a settled outcome.
Controlled tests run inside the company produced clearer examples of the same pattern. Every test was constructed to force the worst plausible response; none reflected behavior observed in normal operation.
Hubinger returned to the same distinction when he addressed the 10 percent figure again. The number applied only to a later stage in which recursive self-improvement had already produced systems far beyond current scale. Present models, he wrote, still carried low risk. The gap between the two regimes remained the central unknown.
Mrinank Sharma left Anthropic in February 2026 after leading the Safeguards Research team. In his departure note he stated simply that the world was in peril.
Alex Turner resigned from Google DeepMind four months later. He cited the laboratory’s acceptance of a Pentagon contract that contained no restrictions on autonomous weapons development.
These exits formed part of a longer sequence of departures from laboratories working at the frontier. The pattern left open the question of what further departures or public statements might follow as the same pressures continued to operate inside the remaining teams.
In early September 2026 Senator Bernie Sanders and Representative Greg Casar introduced the Ban Artificial Superintelligence Act. The measure would permanently prohibit the development and deployment of any AI system that matches or exceeds human cognitive performance across broad domains. It would also bar systems capable of disempowering humanity and would halt advanced AI work until a federal regulator established binding safety requirements.
The bill’s text defined the prohibited threshold in terms of capability rather than intent.
