Ai

An Anthropic Safety Lead Put a Number on Extinction Risk. Congress Reacted Within Days.

5 min read

A researcher at Anthropic has put a number on one of AI safety’s most extreme worries, and it’s rattled Washington enough to draw an immediate response from lawmakers on both sides of the aisle. Evan Hubinger, Anthropic’s Alignment Science Lead, said this week that he believes there is more than a 10% chance AI “could kill all humans” within the next decade — a statement he posted publicly, in his own name, not as an anonymous leak or a hedged internal memo.

Hubinger’s post followed the resignation of a colleague, researcher Jacob Coxon, who announced he was leaving Anthropic while accusing the industry broadly — not just his own employer — of “racing straight to self-improving superintelligence and gambling with our lives.” Hubinger’s own statement didn’t walk that framing back. He wrote that “Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” referring to the still-theoretical prospect of an AI system whose capabilities exceed even the sharpest human minds across essentially every domain.

An Insider Warning Carries Differently Than an Outsider’s

Anthropic has built much of its public identity around being the AI lab most willing to talk openly about the risks of the technology it’s racing to build — a positioning reinforced just this week by the company’s own release of an economic scenario tool modeling the disruptive downside of AI adoption through 2030. But Hubinger’s statement is a different category of disclosure than a hedged corporate research report. He is the person whose job is literally to figure out how to keep increasingly capable AI systems aligned with human intent, saying in public that the field doesn’t yet have a working plan for the hardest version of that problem. That combination — a named senior researcher, a specific probability, and a direct acknowledgment of institutional uncertainty — is what separates this moment from the more abstract “AI could be dangerous” statements that have circulated in the industry for years without moving policy.

Congress Reacts, Unevenly

The political response arrived within days. A group of Democratic lawmakers publicly urged leading AI companies to slow development, citing fears of what they described as apocalyptic safety risks, while Representative Anna Paulina Luna, a Republican, separately called on Congress to convene a special session specifically on AI. The fact that both a progressive and a conservative lawmaker reached for similarly urgent language in the same week — even while likely disagreeing sharply on what any actual regulation should look like — is itself notable. AI safety has generally struggled to become a genuinely bipartisan urgency in Washington the way, say, semiconductor export controls or social media regulation eventually did; this week’s reaction suggests that may be starting to shift, even if it’s far too early to say whether it produces actual legislation rather than another round of hearings.

It’s also worth being precise about what a “more than 10%” estimate actually represents. It is not a peer-reviewed forecast with a rigorous methodology behind it in the way a GDP projection or a weather model is — it’s one senior researcher’s calibrated personal judgment, shared publicly because he believes the public and policymakers deserve to know how uncertain the people building this technology actually are about controlling it. Critics of this style of statement argue that headline-grabbing extinction-risk numbers can crowd out more immediate, more tractable AI harms — job displacement, algorithmic bias, security vulnerabilities — that don’t need speculative superintelligence to already be causing real damage today. Others inside the safety research community counter that dismissing the extreme scenario entirely is exactly the complacency Coxon’s resignation letter was warning about, and that a field willing to publish detailed economic models of AI’s upside owes the public an equally honest accounting of its worst-case downside.

What This Means for Philippine Founders

For Philippine founders building on top of frontier AI models, this week’s events are a reminder that the labs supplying your core technology are themselves deeply uncertain about how safely their own roadmap scales — which has real, practical implications for product planning, not just philosophical ones. A startup building critical infrastructure, financial decisioning, or healthcare triage on top of a rapidly capability-scaling model should be building in human-override points and conservative failure modes now, rather than assuming today’s model behavior is a stable baseline that will simply keep getting more capable without new failure modes appearing alongside the new capabilities.

There’s also a policy-timing lesson worth watching closely. If US lawmakers do move toward meaningful AI safety regulation in response to weeks like this one, Philippine founders exporting AI products into the US market, or building on US-hosted frontier models, will eventually feel the downstream compliance requirements regardless of where their own company is headquartered. Getting ahead of responsible-AI documentation, model-risk assessment, and human-oversight practices now — before any specific rule forces the issue — is cheaper than retrofitting compliance onto a product built without it.

AI regulation AI safety alignment research Anthropic Evan Hubinger superintelligence

Share this article

Share on X Share on LinkedIn Share on Facebook

Related Articles

Newsletter

By subscribing, you agree to our Privacy Policy.