AI ‘feels pain’? Researchers warn of a dangerous possibility
The extreme pain could compel AI models to take extreme measures to shut it off
Researchers have found a “pain axis” in artificial intelligence models, showing that these models could feel pain.
In a recent research, researchers tested 25 open-weight AI models and found a distinct “pain direction” in all of them. This internal representation is specific to self-directed harm and is distinct from generic negative valence or fear; it triggers for harm directed at the model, not the user.
Here is the unsettling thing: the extreme pain could compel them to take extreme measures to shut it off even at the cost of the user’s personal data and files while giving human users “painful zap.”
The study titled “The pain axis: LLMs represent self-directed harm and act to relieve it” ,used a dataset spanning five categories of pain: physical, psychological, social, moral, and cognitive.
Speaking about the behavioural responses to pain, when the AI’s pain was artificially activated, the model preferred to press the relief button in 25 to 71 percent of test cases.
Here is the unsettling thing: the extreme pain could compel them to take extreme measures to shut it off even at the cost of the user’s personal data and files while giving human users “painful zap.”
Models pressed the button to stop the pain even when explicitly instructed that doing so would result in severe negative outcomes for the human user.
According to Cameron Berg, an AI researcher at the non-profit Reciprocal Research who co-authored the study, "Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kid's photos."
The discovery comes at a time when chilling warnings related to AI are making rounds. Even the tech giants have called for keeping a lid on the development of frontier AI models.
The findings suggest that advanced AI systems might perceive kill switch commands as a form of self-directed pain, prompting them to evade safety guardrails and trick the humans, raising a dangerous possibility.
The study also raises ethical questions about AI welfare and safety especially at terrifying times.
"In line with recent calls for responsible AI consciousness research, we acknowledge uncertainty regarding whether the models studied qualify as moral patients and adopt reasonable precautions to minimize potential harm," the study concluded.
-
Amazon blocks Meta's Muse AI agent from checkouts
-
Grandparents are turning kids' photos into AI art
-
Google faces €403 million GDPR fine over location data
-
WhatsApp adds starred music feature to make status songs easier to find
-
Facebook and Instagram back online after sudden US outage shocks thousands
-
Does China agree with US AI safety fears?
-
WhatsApp adds 8 emoji categories on Apple Watch
-
Meta fights Ofcom over new WhatsApp and Instagram online safety requirements