News

Anthropic researcher warns AI has “>10%” chance of killing all humans this decade

Anthropic researcher warns AI has “>10%” chance of killing all humans this decade

Talk of AI “wiping out humanity” used to sound like science fiction or niche research warnings.


In September 2026, a researcher at the center of one of the world’s leading AI labs put a concrete number on that risk in public.


Evan Hubinger, Alignment Science Lead at Anthropic, said on X (formerly Twitter) that he personally estimates a greater than 10% chance AI could kill all humans within the next decade.


Outlets including BBC, CBS, and Forbes covered the post, and it spread widely in a short time.


Because sensational headlines travel fast, it is worth separating what was said from what was not.


What happened


The exchange began when Anthropic researcher Jacob Coxon announced his resignation.


Coxon said people building frontier AI earnestly believe it could kill us all by the end of the decade, and criticized OpenAI and Anthropic for not acting responsibly.


Hubinger replied that Coxon was right, saying they really do earnestly believe AI could kill all humans.


He then gave his personal estimate: more than 10% within the next decade.


He also said he believes Anthropic is trying its best, but the company does not yet have a plan to solve alignment for superintelligence and is not clearly on track.


A safety lead acknowledging in public that the plan is still missing is unusually blunt for the industry.


This is not about today’s models being dangerous


Hubinger also said the risk from currently existing models is low.


His main concern is future superintelligence—especially recursive self-improvement.


The backdrop is a sense that AI improving itself beyond human pace may be arriving faster than expected.


In other words, this is not “today’s Claude or ChatGPT will end humanity,” but “we still lack a way to keep future superintelligence aligned with human intent.”


Alignment means ensuring increasingly powerful systems reliably pursue goals compatible with human intentions and safety.


Capability races are ahead of control design—that is the core of the warning.


Why “10%” lands hard


In everyday probability talk, 10% can sound like “sometimes.”


When the subject is the survival of humanity as a species, the same number reads very differently.


Few people would board a plane after being told there is a 10% chance of a crash.


A senior researcher at a major AI lab stating that level of estimate in public is itself the news.


This remains a personal estimate, not an official company position.


Still, hearing an alignment lead say both “above 10%” and “we do not yet have a plan” shows a temperature that safety marketing alone cannot capture.


It also fits a wider industry mood


Since around 2023, leaders at OpenAI, Google DeepMind, and Anthropic have also spoken about existential AI risk.


More recently, pauses in developing models with advanced cyber capabilities and stronger third-party evaluations show more “stop before shipping” behavior.


At the same time, labs cannot easily exit the next-model race, so internal tension continues: people know the danger and still push forward.


Coxon’s resignation remarks and Hubinger’s reply made that tension visible outside the company.


The important point is not scare theater, but that people building the systems are saying they earnestly believe the risk is real.


Takeaways for businesses and developers


Extinction-risk debate can feel distant, but the operational implications are concrete.



  • Keep agent and tool permissions minimal by design

  • Put human final approval before self-improvement, auto-deploy, or autonomous actions

  • Re-run safety evaluation and monitoring after every model update

  • Design systems that can be stopped—do not connect everything just because it is convenient

  • Keep watching vendor safety policies, evaluations, and incident reports


Even for smaller companies, basics such as stronger public-API authentication, checking exposed admin screens, and log monitoring remain effective.


Superintelligence debates and today’s operations sit on different layers, but thicker permission design and human checkpoints are shared culture.


Summary


Anthropic’s Evan Hubinger personally estimated more than a 10% chance that AI could kill all humans within the next decade.


He said current-model risk is low, while warning about superintelligence from recursive self-improvement and the lack of a clear alignment plan.


Beyond the shocking headline, the heavier practical message is “we do not yet have a plan” and “we are not clearly on track.”


At Makoto Tejima, when planning AI agents and business automation, we emphasize design that includes existing-system integration, permission management, monitoring, and human confirmation points.


Please feel free to consult us about secure agent design, hardening of public services, or adding AI features to existing systems.