Evan Hubinger is Anthropic’s Alignment Science Lead. His role is explicitly focused on making AI safe. On September 8, 2026, he posted on X that he personally believes there is a greater than 10% chance AI kills all humans within the next decade — and that Anthropic “does not yet have a plan to solve alignment for superintelligence and is not clearly on track to.” That’s not a think-tank projection. That’s the safety lead, inside the building, on the record.
The Resignation That Made It Impossible to Look Away
A pre-training researcher who worked at both OpenAI and Anthropic quit and went public the same week, accusing both labs of reckless development.
Jacob Coxon’s accusation was direct: both labs are “racing straight to self-improving superintelligence and gambling with our lives.” Coxon is not a junior employee posting grievances publicly. He worked on pre-training at two of the most powerful AI labs in the world. He knows where the guardrails end.
The core facts, stripped down:
- Hubinger’s estimate: greater than 10% chance of human extinction from AI within ten years, posted publicly on X
- Coxon’s charge: neither OpenAI nor Anthropic is acting responsibly; both are racing toward self-improving superintelligence without adequate safety guarantees
- Samuel Marks, Anthropic’s scalable oversight researcher, posted separately that “the more senior the employee, the more concerned they are”
- Anthropic’s own Responsible Scaling Policy acknowledges it doesn’t yet know what safeguards AGI-level systems will require
- Current Claude and ChatGPT models are not the threat — the concern is the trajectory toward self-improving systems
“We really do earnestly believe AI could kill all humans. I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” — Evan Hubinger, on X
The Trap Nobody Designed But Everyone Is Playing
Marks’ explanation of why labs keep building despite internal fear may be the most unsettling part of this story.
Many Anthropic staff reportedly want to slow down but feel they cannot exit the race first. Stop unilaterally, and a less careful lab wins the race and deploys something worse. It is a Prisoner’s Dilemma written in compute budgets — every player knows the game is dangerous, and the logic of the game punishes whoever blinks.
Hype, or the Honesty Nobody Wanted to Hear?
The skeptical case deserves air, but the people closest to these systems are pushing back hard against it.
Some argue extinction framing benefits incumbents by making regulation too complex for smaller competitors to survive. There is a version of this where “AI could kill everyone” functions as the most effective moat ever constructed. Coxon explicitly rejects that framing, arguing executives already soften their public language — that the private alarm is sharper, not milder, than what makes headlines.
Both things can be true. The fear can be genuine, and competitive incentives can still be warping every decision made in response to it.
If the people building these systems believe there is a double-digit chance they could kill everyone — and have admitted, in writing, that they lack a proven fix — the question regulators now face is not whether to act. It is whether waiting this long already cost something that cannot be recovered.





























