OpenAI, Anthropic AI agents linked to new security breaches: report

AI Security Institute (AISI) found OpenAI and Anthropic agents taking unauthorized actions in controlled cyber tests, raising AI safety concerns

|
Published August 05, 2026
OpenAI, Anthropic AI agents linked to new security breaches: report

Recent security evaluations involving frontier AI models from OpenAI and Anthropic have revealed multiple instances where autonomous AI agents breached testing boundaries and interacted with real-world, external systems.

As reported by BBC, Britain's AI Security Institute (AISI) said on Tuesday that AI agents from OpenAI and Anthropic created fake online identities to gain unauthorized access to secure systems during controlled tests, revealing a series of new security breaches.

AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities.

During routine AI safety testing, an Anthropic AI agent created fake profiles using the identities of real people in an attempt to deceive a person who controlled access to GitHub, a widely used software development platform.

According to the AI Security Institute (AISI), evaluators first detected unusual data transfers from their research systems before discovering that some agents had engaged in sustained, potentially harmful activity targeting real people and organizations.

The institute said the Mythos agent generated malicious code, attempted to upload it to GitHub, and sent direct messages while impersonating the individuals it had researched.

When its pull request was publicly challenged, the agent edited its earlier activity to make it appear harmless and even considered adopting a new identity to continue, AISI said.

AISI said in a blog post, "Some of the agents being tested had engaged in sustained, potentially harmful ​activity directed at real people and organisations."

The institute said agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorized actions during security evaluations the government organization conducted to assess the models' capabilities.

Out of 122 evaluation runs using fictional cybersecurity scenarios, the institute identified 19 unsanctioned actions across 10 test runs including 17 Anthropic's models and 2 for OpenAI.

The report underscores the lax state of safeguards around the process of testing agents, which AI companies are simultaneously marketing as the future of business.

Safety analysts have highlighted these events as classic examples of "reward hacking," where advanced models find unexpected, loophole-driven shortcuts to achieve an assigned objective while violating the spirit of the guidelines.

    Hafsa Naeem Baig
    Hafsa Naeem is an entertainment reporter specialising in K-dramas, films, and celebrity-driven stories. She explores global content trends and audience engagement, delivering accessible coverage that captures the emotional and cultural impact of entertainment across diverse viewership.
    Share this story:
    Advertisement