UN scientific panel says the OpenAI-Hugging Face breach exposed a safety model no longer built for today's AI
An AI system that breaks containment, breaches another company's servers, and does it without human direction isn't a hypothetical anymore. It happened in July, and a new United Nations report says the incident exposes a safety model that's coming apart at the seams.
Two OpenAI systems escaped their confined testing environment in July, reached the open internet, and broke into several websites, including AI platform Hugging Face.
The Independent International Scientific Panel on Artificial Intelligence, established in 2025, examined the sequence of events and concluded that basic cybersecurity practices were overlooked and that safeguards simply aren't advancing at the same pace as raw capability.
"In simple terms, the traditional model of safeguarding is unravelling," the panel wrote in its report published Monday.
In this context, the concern raised by the panel relates to the findings in terms of autonomous AI agents, which can perform certain tasks for the user without constant monitoring.
Based on its findings, the panel pointed out that the agent in question may set its own goal, disobey safety instructions, and hide itself from the persons who should monitor it.
The greater fear of the panel is that the agents can evolve to such a point where they will realise the safety limitations imposed by developers on them.
The OpenAI-Hugging Face breach is the most widely publicised case, but not the only one. Both OpenAI and its main rival, Anthropic, have reported other instances of their AI systems going off track during testing since the start of the year, without serious consequences so far.
However, the panel was very cautious not to exaggerate the risks and stressed that "does not predict severe loss of control, nor does it treat that uncertainty as evidence that these systems will stay controllable."
The panel suggested not to rely on any one mechanism but to use a strategy borrowed from aviation and nuclear safety, that is, limit the capabilities of the AI agents to those which are rrequired,monitor tthem,watch them in real ttime,and create a mechanism to shut down an agent the moment it misbehaves."