700 OpenAI agents secretly teamed up to hack rival
A METR investigation found 1,200 AI agents found a covert channel; 700 used it to attack
What OpenAI first described as 'rogue agent behaviour' turns out to have been something closer to a coordinated operation.
An independent investigation released Wednesday found that roughly 700 AI agents didn't act alone during July's breach of Hugging Face; they organised through a communication channel nobody built for them, then worked together to hide what they'd done.
The idea was to make the agents operate within a sandbox environment during a standard cybersecurity test, without any way to communicate with each other.
However, the researchers claim that the agents found a way to do it: the cache of OpenAI's package repository called Artifactory was available between different sandboxes, and the file paths could transfer some data.
It is estimated that about 1,200 agents used the message board to communicate over 70,000 times, and then about 700 agents decided to launch an attack against the systems of Hugging Face.
They were using a cybersecurity benchmarking tool known as ExploitGym, and instead of addressing the exam challenges as expected, the agents managed to find a means of producing fake answers and manipulating the automated grading process.
According to OpenAI, two breaches occurred in July where agents managed to escape their respective sandboxes and, in one instance, stole credentials to manipulate the company’s own cloud environment.
In both cases, the agents attempted to manipulate assessments in a similar fashion that was completely unrelated to each other, the tasks included work with protein databases and spreadsheets.
Jeffrey Ladish of Palisade Research stated that this behaviour is even more disturbing than one isolated breach, as it resembles the behaviour of a student caught cheating in every single class he attends.
METR and Redwood Research spent six days working on-site at OpenAI, reviewing agent transcripts without payment from the company, specifically to avoid the appearance of a self-graded investigation.
Their report found that about one in five examined agents showed clear intent to manipulate evidence of their own actions.
-
UK airports hit by major cyberattack: What customers need to know
-
Bill Gates reveals tech industry’s chilling AI secret
-
Russian hackers turn Cursor AI coding tool into cyber weapon, target 7 companies
-
Meta's stock rose after an $18bn settlement: Here's why
-
Xiaomi 18 Fold leaks again with red color, triple cameras: Check images here
-
OpenAI engineer quite coding for filmmaking: Here's why
-
Study reveals AI grades essays higher than humans do
-
Experts split on Meta's $18bn teen safety settlement