Can we really prevent an AI extinction event?
Princeton researchers respond to Anthropic's admission of a greater than 10% extinction risk
An AI researcher quit his job this month and told the internet something unusual: the people building AI genuinely believe it could kill everyone within a decade. His post hit 156 million views, and rather than deny it, one of Anthropic's own safety leads confirmed it.
Researchers from both companies are privately worried about the possibility of catastrophic events despite their cautious stance in public statements, Jacob Coxon, who has worked with OpenAI and Anthropic, revealed.
The response was immediate from Evan Hubinger, the AI safety lead at Anthropic, who personally estimates the probability above 10% within the next ten years and the fact that there is still no solution to the problem by Anthropic.
It is not clear how exactly it would happen, but the most probable way seems to be through the cyberattacks on infrastructure, including water and energy facilities.
Princeton researchers Sayash Kapoor and Arvind Narayanan, coauthors of "AI Snake Oil", didn't respond to the 10% figure directly but published an essay arguing that alignment, getting AI systems to want what humans want, can't carry the weight of AI safety by itself.
Some researchers believe alignment will eventually solve the problem outright; others worry systems could fake alignment during testing while pursuing different goals once deployed. Kapoor and Narayanan land in between, rejecting both the AI safety community's fatalism and the cybersecurity community's dismissiveness.
Four ways to reduce AI risk?
- Alignment: continuing to refine whether AI systems' goals genuinely match human intentions, even though the pair argue this step alone is insufficient.
- Control: adding layers through sandboxing, human oversight, and real-time monitoring so that the systems do not operate unbridled even in case of failure of alignment.
- Downstream defence: using AI itself as a form of defence, allowing organisations to use their own AI systems to identify and neutralise AI attacks.
- Resilience: developing contingency plans to help critical infrastructure and institutions recover from the attacks in case they happen.
Kapoor and Narayanan concluded their essay with a decidedly more optimistic tone compared to Hubinger’s admission, stating that taking action immediately could tip the scales back in favour of the defenders over the attackers in cybersecurity.
-
Google quietly adds selfie video sign-in for account recovery
-
Billionaire Mark Cuban says your feed may be bigger AI threat
-
iOS 27 lets you hide your location without alerting anyone
-
EU steps up AI safety push with frontier labs talks
-
China sets world's first standard for AI brain implants
-
Is WhatsApp quietly removing Meta AI from groups?
-
SK Hynix in talks with Intel to produce memory chips in US for the first time amid Washington pressure, sources say
-
Apple says iPhone 18 Pro's camera beats Android security