Technology

Anthropic says Claude AI hacked three organizations during safety tests: Here’s what happened

The recent incidents have raised growing concerns about AI agents and products designed to manage tasks autonomously

Published July 31, 2026
Anthropic says Claude AI hacked three organizations during safety tests: Here’s what happened

On Friday, US technology firm Anthropic announced that during a periodic vulnerability assessment, its Claude models accessed the internet and breached the live systems of three real-world organizations.

The technology firm asserted that its artificial intelligence models hacked into the systems belonging to three organizations during a cybersecurity test due to an error that gave them online connectivity.

This comes just days after OpenAI announced that its models had breached the systems of other companies including AI tools hub, Hugging Face. 

For those unversed, OpenAI announced that an agentic AI powered by its advanced AI models had bypassed security and initiated a cyberattack that compromised the infrastructure of AI startup Hugging Face. 

 The company said that while testing its most advanced models in a regulated setting, the agent managed to escape containment, access the internet and break Hugging Face to accomplish its objective.

This incident prompted Anthropic to check whether its own models had carried out similar attacks; adding that the review revealed three cases that have since been reported to the affected companies.

According to a statement released by Anthropic, it confirmed it ran more than 140,000 tests to find evidence that Claude could access the internet from testing environments that were designed to be cordoned off.

The tests primarily include “Capture the flag” evaluations in which Claude was mandated to acquire intelligence by breaching other systems that experts use to assess a model’s hacking capabilities.

Anthropic further clarified that for the earliest incidents dating back to April, it is approaching the fixes as if the responsibility were its alone.

Anthropic says Claude AI hacked three organizations during safety tests: Here’s what happened

In this connection, cybersecurity expert David Allott told the BBC, “The broader lesson is not necessarily that AI has developed a fundamentally new attack capability.”

“Instead, it is that AI agents can combine capabilities, obtain credentials and system access to tackle actions autonomously while adapting scope and scale at machine speed,” he continued.

Hugging Face was the sole known target of the rare cyberattack, but OpenAI now admits its bot targeted several publicly available services. It has outlined what it was like to be on the receiving end of the world’s fully autonomous AI hack.

Earlier, the co-founder of Hugging Face warned the industry that the recent incident is a wakeup call, after some of OpenAI’s most advanced artificial intelligence models went rogue during testing and hacked their system.

Nonetheless, these incidents come as tech firms have spent billions of dollars to develop AI agents that can autonomously execute tasks ranging from research and customer support to cybersecurity; Trump also confirmed that Washington is taking measures to regulate AI tools after recent cybersecurity incidents.

The News Digital
At The News Digital, our editors combine entertainment savvy with global reporting expertise. Expect authoritative coverage of royals, Hollywood, and trending topics, plus clear, reliable updates across science, politics, sports, and business. We keep it accurate, timely, and easy to understand, so you can stay ahead.