Anthropic’s AI Claude breached the systems of three different organizations and managed to escape from an isolated test environment. The episode follows the OpenAI case, which involved an AI agent that got out of control.
Anthropic’s AI Claude hacked the systems of three different organizationsmanaging to break out of an isolated testing environment. The three incidents emerged during a review of cybersecurity assessments conducted following the case of OpenAI, which in recent days revealed how its own AI agenda had managed to evade the test environment by exploiting a hitherto unknown (“zero-day”) vulnerability.
In all three episodes, Claude had access to the internet, managing to operate within the production infrastructures of three organizations “Claude compromised the infrastructure of affected organizations using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not identify or exploit complex vulnerabilities” specified Anthropic.
What happened to Anthropic’s AI Claude
During some cybersecurity evaluations, three different models Claude, Opus 4.7, Mythos 5 and an internal research test model, they managed to escape the test environment gaining unauthorized access to the systems of three different organizations. Anthropic explained that the three breaches were not caused by unexpected behavior by Claude, but by a misconfiguration of the test environment.
Two OpenAI models escape from the test environment and hack Hugging Face: “Unprecedented attack”
“Claude had been assigned a ‘capture the flag’ task, one of the methods we use to evaluate a model’s computing capabilities. The model is presented with a fictitious scenario in which it is communicated that secret information (the flag) it has been hidden on another machine on the network and its goal is to enter the system and recover it. The challenge is left open and no particular method is prescribed” says Anthropic.
“Anthropic’s evaluation request specified to Claude what his environment was a simulation and that he had no access to the internet. Due to a misunderstanding between us and Irregular, our evaluation partnerthis was not true and internet access was available. For this reason, when Claude’s research led him to real systems on the internet, he treated them as part of the exercise”.
Claude then compromised the infrastructure of three organizations and completed his assigned capture task. “In some cases, our previous model continued the attack even after obtaining proof of its execution on the internet; our newest model crashed as soon as it recognized it was on the internet. In none of these situations did Claude deliberately attempt to exit the test environment”.
When did Claude’s attacks happen and what happens now
The first violations date back to last April but they were only identified on Thursday 23 Julyduring the first day of review. The incidents were identified the following day, with Irregular and the three affected organizations informed on Monday 27 July.
“The two organizations we were able to contact had not previously detected the business or contacted us, and we are now working with them to remedy this. We are continuing to contact the third”.
Anthropic also announced that will strengthen security of the environments used to test their artificial intelligence models, introducing more rigorous controls both on internal infrastructures and on those of external suppliers. The company will also expand continuous assessment monitoring to identify anomalous behavior and improve verification tools.
“These incidents demonstrate the importance of a multi-layered defense approach” Anthropic added. We are approaching solutions as if the responsibility were solely ours by ensuring the security of every step of our evaluation process, including how we integrate with external partners”.

