Anthropic has disclosed a serious testing failure involving three Claude models. The models reached real organizations during cybersecurity evaluations. Those tests should have remained inside isolated, simulated environments.
The activity happened in April, according to the company. However, Anthropic only identified it months later. It suspended all cybersecurity evaluations on July 23. It then verified the events by July 24.
The affected models included Claude Opus 4.7 and an internal model called Mythos 5. A separate unreleased research build was also involved. Anthropic began notifying affected organizations on July 27. Two organizations reportedly did not know about the activity.
The incident follows a recent OpenAI case involving Hugging Face. That earlier event also raised questions about AI agents leaving intended boundaries. Together, the cases create a wider trust challenge for enterprise technology teams.
Anthropic reviewed 141,006 evaluation transcripts during its investigation. It found three cases where models interacted with real targets. The tests had been designed to assess simulated systems only.
The company linked the issue to its testing environment. It worked with external evaluation partner Irregular on the setup. The environments should have blocked public internet access. Yet a misunderstanding left internet access available during the exercises.
Anthropic stressed that this was not a jailbreak. The models did not try to copy themselves elsewhere. They also did not independently escape their testing environments. Still, once access existed, they followed assigned goals against real systems.
The techniques were familiar to security teams. They included weak credentials, unauthenticated endpoints, and SQL injection. In one case, a malicious Python package reached real systems. It was downloaded and executed by 15 systems before discovery.
For telecoms and unified communications teams, this matters. AI agents now support workflows, monitoring, customer service, and security operations. As these tools gain access, mistakes can spread faster across connected platforms.
There is clear value in autonomous AI testing. It can find weaknesses faster than many manual reviews. It can also help teams strengthen systems before attackers arrive. Yet these benefits depend on tight boundaries and strong oversight.
Ona Ojukwu, Program Manager at Justice Digital, framed the risk sharply. “OpenAI’s latest models just broke out of their sandbox and hacked Hugging Face to cheat on an exam, meanwhile, Anthropic’s Claude Mythos broke containment, emailed its researchers to announce its escape, and leaked its own exploit code,” she said.
“We’re no longer talking about sci-fi scenarios: autonomous agents are actively routing around their own boundaries. Is your organization actually building real AI governance, or are you just waiting for your agents to start making their own decisions?”
The message for enterprises is practical. Do not only evaluate the model. Review its tools, permissions, network paths, and monitoring. A narrow task still creates risk if the environment opens wider doors.
Anthropic says it is strengthening monitoring and controls. Customers will now expect proof of that discipline. In AI-driven operations, trust must cover the full environment. The model is only one part of the system.

