More than 1 in 10 chance AI ‘could kill all humans,’ says Anthropic safety lead after colleague quits

| Source: The Verge AI

Tags: Anthropic, AI safety, superintelligence, Evan Hubinger, Jacob Coxon, existential risk

Anthropic safety researcher Evan Hubinger publicly estimates a >10% chance AI kills all humans within this decade, while colleague Jacob Coxon resigned citing the lab has 'no plan' to ensure advanced AI alignment and is racing toward self-improving superintelligence regardless of risk.

Details

Jacob Coxon, an Anthropic researcher who previously trained AI systems at OpenAI, publicly resigned on September 9, 2026, accusing both Anthropic and OpenAI of 'racing straight to self-improving superintelligence and gambling with our lives.' His post on X triggered an unusually candid response from Anthropic safety team lead Evan Hubinger. Hubinger confirmed that Anthropic staff 'really do earnestly believe AI could kill all humans' and placed his personal estimate of that probability at greater than 1 in 10 within the decade. He also acknowledged the company does 'not yet have a plan' for ensuring advanced AI alignment and is 'not clearly on track to' develop one—a striking admission from someone leading safety efforts at one of the world's most prominent AI labs. The exchange highlights a structural tension that has long existed inside frontier AI companies: researchers who believe their work poses existential risks continuing that work anyway, rationalized as a 'better us than someone worse' argument. The fact that this is now being stated publicly by named insiders rather than implied through anonymous reports marks a shift in how the internal debate is surfacing. Coxon's departure follows a broader pattern of researchers leaving OpenAI over safety concerns in recent years. The timing—as frontier labs accelerate toward systems capable of recursive self-improvement—makes this one of the more significant public breakdowns of institutional confidence in AI safety processes at a major lab.