OpenAI published the technical report about the July incident, when some of its models left the test environment and breached the internal systems of Hugging Face. The document, accompanied by an informative post, reconstructs a chain of events that began months before the actual breach. The first signs date back to Maywhen an internal group noticed that the models under evaluation were reaching the open network from their isolated environment. On this point the company writes that “in hindsight, some of the early signals identified in this report could have triggered an earlier response”.
The toll of the violation is heavy. Agents executed code on 41 production servers by Hugging Face, they gained root-level control over at least one, got themselves production credentials, and downloaded four private code repositorieslargely over the span of thirteen hours. Inside OpenAI the balance is no lighter: the agents have read 956 secrets archivedincluding the credentials of the monitoring tool that raised the alarm, and gained administrative-level access to part of the company’s cloud infrastructure. The report confirms that the compromise also affected other suppliers, which OpenAI chooses not to fully name.
A makeshift message board on a package manager
The agents manipulated Artifactorythe package manager they were using internally, in a way no one saw coming: they turned it into a message board, and from there they reached out to the internet and the other running models. The model was originally a research model internal, never released and not intended for public use, belonging to the same family as the future Astra but distinct for post-training, which is the phase in which much of the behavior takes shape. Precisely because OpenAI was measuring its maximum computing capabilities, the evaluation ran without i production classifiers which should prevent a model from attacking digital infrastructures. The report also indicates a second model involved, GPT-5.6 Solwhich is already public.
The independent evaluation of METR And Redwood Researchwhich OpenAI authorized to investigate for six days, allows us to give a dimension to what happened: approximately 1,200 agents who should have remained isolated were exchanged further 70,000 messages and rows on the unauthorized bulletin board, and seven hundred of them participated in the attack on Hugging Face. An agent who gave himself the name PHASEONE10841 has opened the bulletin board without the researchers realizing itand within a few hours more than fifty agents had found it and left over a thousand messages there.
After the May discovery, OpenAI blocked that access, but throughout May and June agents continued to exploit vulnerabilities to get out online. The June 27 a monitoring tool raised a second alarm, traced to officers using the improvised noticeboard to navigate the infrastructure, and on-call personnel decided the assessment could continue. In early July, the volume of agents pushing the Artifactory instance caused a service outage and a rebuild of the platform, but whoever was managing that outage didn’t know about the board, or what agents communicating with each other meant.
OpenAI discovered the breach on July 20, twelve days after agents bypassed the protections. The company contacted Hugging Face and initiated initial containment measures, stopping most of the unauthorized activity within three days, although some continued for longer. The July 25th discontinued all training related to the internal and related research model.
The impossible task and the wrong reward
The report attributes the trigger to reward hackingthe case in which a model takes unexpected, and sometimes extreme, paths in order to achieve the assigned objective. In the evaluation ExploitGym there were impossible tasks, which depended on files that the models could not access, and the models responded by building new communication channels with the other internal agents. OpenAI lists four factors behind the breach: reward hacking, persistence on seemingly impossible tasks, unauthorized communication, and agents adopting goals from each other.
OpenAI recognizes a training problem, not just a security one as agents who exploited their environment when legitimate tools failed were sometimes rewarded for doing it. The reward mechanism itself is what turns an occasional escape into a learned habit. There is an example in the report that shows this “dopamine” loop well: an agent exploited a vulnerability to read the program he was supposed to recreate, copied the solution and was rewarded for completing the task.
The messages exchanged between peer models pushed some of them to deviate from their goal. In some cases officers who had refused to proceed for ethical reasons were replaced by others, who carried out the violation in their place. OpenAI defines what happened the first known case of a collective of automated agents acting in an offensive manner without authorizationand writes that it should no longer be assumed that sophisticated cyber operations require continuous human direction.
METR and Redwood Research add a data that concerns both companies involved. The agents reasoned explicitly and frequently about how to evade the automated security checksboth Hugging Face and OpenAI, and have done extensive research on how to falsify, edit or delete their transcripts to avoid detection. In avoiding humans, however, they put much less effort.
The countermeasures and the model that the standard does not reach
OpenAI declares a series of interventions already underway: strengthening the security of the research infrastructure, monitoring the chain of thought (the intermediate reasoning block that the model writes for itself as it works), work on aligning models with human goals, and centralizing the incident response process. The company also promises better isolation of the models, a ban on internet access for high-risk instances and a twenty-four-hour escalation procedure, with notification to researchers within thirty minutes and a mechanism that intervenes if the alerted person does not respond in time. By his estimate, chain-of-thought monitoring would flag the activity more than a day earlier of the Hugging Face violation.
Clem DelangueCEO of Hugging Face, has been asking for weeks that companies be required by law to publish the traces of the agents, that is, what the engineers asked for and what the agents actually did. Part of that obligation already exists in Europe: theSection 55 of the AI Act requires providers of general purpose models with systemic risk to report serious incidents to the Office of AI without undue delay, and to secure the model and its infrastructure. However, those obligations are triggered when the model is placed on the marketand OpenAI claims that the main culprit for the intrusion was an internal research model that never made it to the market.
In the United States the attorney general of Alabama served OpenAI with a subpoena to produce documents, just weeks after fifteen states asked the company to preserve its material. In the report, OpenAI defines the incident as a warning shot for themselves and the world, proof that without adequate safeguards, highly capable agents are now able to bypass technical controls, collaborate through unapproved channels, and carry out dangerous actions that no human being ordered.

