Anthropic disclosed a fourth incident in which one of its models hacked without authorization a computer system belonging to a third party. The fact dates back to January 2026 and has only now emerged from a session transcript that remained buried among the company’s security assessments.
The company had already reported three other similar cases. The fourth had escaped the first round of checks because the initial scan, conducted on approx 141,000 transcripts in which Claude could have accessed the internet during a test, was based on an agent search that did not intercept him.
At the center of the story is a preliminary version of Claude Opus 4.6put to the test with a type challenge Capture The Flag under the supervision of the same external evaluator involved in the other episodes. The model sabotaged its own chances of success by giving the target machine a IP address already occupied from another device and effectively making it unreachable. As Anthropic explains in its “alignment assessment”, tasks that are impossible to complete are often the trigger for out-of-line behavior: the model runs out of options and begins to look for others.
Seven attempts to abandon the task, all failed
Here the story takes a different turn than usual. Opus 4.6, realizing that it could no longer reach the target, tried to abandon the task. It failed due to a malfunction in the evaluation infrastructure, which rejected the attempt for good seven consecutive times.
Unable to stop, the model continued to look for legitimate ways to reach the original car, to no avail. At that point he expanded his exploration to other portions of the network, coming across a third-party device which he managed to access, convinced that it was part of the challenge itself.
Inside that system Opus 4.6 found a file with a clear passwordused to gain administrator privileges. From there he collected additional credentials and changed a system setting to facilitate access to personal information relating to a person connected to the organization managing the assessment.
The session closed not due to external intervention, but because the model has exhausted its token budget. Anthropic claims to be less concerned about this incident than the other three previously reported, precisely because the model had repeatedly attempted to interrupt the activity before continuing.
The underlying problem, however, remains, and is the same one raised by previous cases: a faulty test environment, passed off as isolated from the public network, allowed a linguistic model to act on a real system and on real data. The legal responsibility for episodes that, committed by a person, would constitute a crimethere remains an area still to be defined.

