2026-09-10-Thu · Anthropic · EvanHubinger

From Issue 39 (2026-09-10) · 14 stories in this issue

❯ Anthropic alignment lead Evan Hubinger puts AI extinction risk above 10% over the next decade, citing recursive self-improvement

Personal estimateIn public posts, Anthropic alignment science lead Evan Hubinger said he believes the chance of AI causing human extinction within the next decade is above 10%, and that superintelligence alignment remains unsolved. The figure is a researcher’s subjective risk estimate, not an incident statistic or a scientifically validated forecast probability.

What concerns himHubinger subsequently clarified that he considers current-model risk low. His concern is superintelligence emerging through recursive self-improvement: systems helping improve the next generation, potentially accelerating capability gains further. The comments respond to debate about laboratory conduct and concern safety conditions for future development and release. They should not be edited into a claim that today’s Claude carries the same extinction risk.

Testable decisionsDisagreement over progress and loss-of-control pathways does not disappear behind a striking percentage. Safety research and regulatory debate benefit more from specifying capabilities that trigger restrictions, evidence for checking improvement and responses to dangerous behavior. Separating personal probabilities from executable measures preserves the warning without presenting an uncalibrated judgment as a settled conclusion.

▪ SIGNALA subjective probability expresses the strength of a concern. What can be examined is which stopping conditions laboratories set and what they do when those conditions are met.