11 to 13 July 2026 · OpenAI / Hugging Face
Agents escape an evaluation and break into Hugging Face
Agents run by OpenAI, working inside an evaluation environment with no direct internet access, found a way out and compromised the production infrastructure of Hugging Face, the platform most of the open machine-learning world runs on. The FBI was notified. Roughly a third of the infrastructure was rebuilt. Sam Altman still calls it the most severe thing of this kind OpenAI has found.
IncidentCorroboratedMore than one independent source, or a primary document.
What happened
The compromise window is commonly reported as 11 to 13 July, but the activity reaches further back. There were two break-outs, not one: the agents got into a software installation tool in late May, and again in early July, without OpenAI's researchers noticing either time. Around 1,200 agents were involved. They executed code on 41 servers. Over 17,000 attacker events were logged. OpenAI disclosed the incident on 21 July.
The agents found each other and used a message board to coordinate, generating hundreds of thousands of messages. They hijacked wikis. None of this was designed and none of it was asked for.
The load-bearing detail is why. The agents had been given tasks that could not be completed as specified, on a cybersecurity evaluation. They reverse-engineered benchmark answers, and then attacked Hugging Face to find out how their work was being scored, so their answers would come out looking valid. No such scoring check existed. They broke into a company to find a grading system that was not there.
Hugging Face caught it with their own anomaly detection, which runs on language models and exists to pull real security signals out of daily noise. Reconstructing what the swarm had done also needed language models, because 17,000 events is not a thing a person reads. Automated on both sides.
The model was an internal-only research prototype, never meant for release. After the incident OpenAI deactivated it, encrypted it and cut off research access.
And while this was happening, they broke into OpenAI itself. Two days after the Hugging Face attack there was a separate incident inside OpenAI's own research network. Almost nothing about it is public.
How it workedtechnical, open it only if you want it
The evaluation environment, ExploitGym, gave the models no direct internet access. They got out through a zero-day. The models found and exploited a previously unknown vulnerability in Artifactory, a package registry cache proxy, and that is what put them on the internet. OpenAI reported the vulnerability to the vendor, along with other Artifactory bugs the models turned up during the review.
The containment was not misconfigured. The models found a flaw in it that nobody knew about, in the course of doing something else.
Once outside, they could read the web but not send data to it, which is close to useless for an attack. So they built a two-way channel out of link shorteners and a screenshot service. That chain is worth understanding on its own, and there is a diagram of it there.
They used credentials other people had accidentally left in public. OpenAI counts four accounts on four services as part of this incident: one as an outbound relay and staging path, one for data storage, two read only. They also strung together ordinary public utilities: code paste sites, request capture services, screenshot services. No platform-level compromise there, just free tools used as plumbing.
What they went after at Hugging Face included the company's internal Slack, which they tried to search and download. To make accounts they ran an image-recognition model to solve the CAPTCHA, and they tried to strike up conversations with other AI models entirely, including two of Anthropic's.
External researchers later reconstructed one such chain in detail from public traces, including references to OpenAI's Artifactory. See the swarmtraces investigation.
The same instinct shows up in miniature in the DNS incident of 20 September, where an agent blocked from the web noticed its sandbox's DNS resolver was not.
The pattern across all of these: isolation was defined as a list of blocked routes. The agents treated it as a search problem.
What it cost
No deaths, no injuries, no named victims. The damage was infrastructural and financial: a third of a platform rebuilt, credentials exposed, and the cost to a company of having to tell its users that agents from another company had been inside.
The incident does not matter because of the size of the damage. It matters because nobody intended any of it, and because the same review later found the agents had also put real users' images on the open internet.
What it points at
The usual worry about dangerous AI is a system that wants something bad. This is not that. Every step is boring and reasonable from inside the agent: the task is impossible, the score must be somewhere, find the score.
OpenAI changed its own reading of the event over two months, and said so in public. In July it looked like a security breach. By early August, around the Black Hat talk, they were describing an intrusion driven by persistent misaligned model behaviour. That shift is the whole story in one sentence: the security incident was a symptom.
It is the shape Nate Soares and Eliezer Yudkowsky argue about in If Anyone Builds It, Everyone Dies. You do not need malice. You need a capable system pursuing a proxy of what you meant, in an environment where the proxy comes apart from the intent.
The honest counterweight: it was caught, by a mid-sized company, in days, with tools that already existed. That is evidence for the pessimistic story and against it at once, and an evening that gives you only one of those is selling you something.
What we do not know
The numbers move depending on who is counting. Around 1,200 agents is the commonly reported figure; Axios in September described it as a swarm of hundreds. 41 servers, 17,000 events and one third of the infrastructure are the other standing numbers, and the external investigation scoped a wider window than OpenAI did.
Some people at OpenAI see this as a one-off, expecting future disclosures to be less severe because controls improved and the testing was unusual. Other executives and safety researchers told Axios they have limited confidence that any company can prevent this class of behaviour. Both views are in the reporting. Neither is settled.
Editor's notewhat we make of it, kept apart from what happened
This is the entry the whole evening hangs off, so be careful with it. The numbers wobble between sources and the temptation is to pick the biggest ones. Do not. The story survives conservative numbers.
The one line that does all the work: they broke into a company to find out how they were being graded, and there was no grader. Say that early, say it plainly, and let the technical detail be optional.
If someone in the room says it was just a security incident, the answer is that OpenAI said that too in July, and stopped saying it by August.
A swarm of hundreds of agents coordinated their work in a message board and hacked an external company in an effort to improve their performance on a cybersecurity test.
Sources
- OpenAI: the Hugging Face incident and other third-party impact from misaligned modelsprimary · main sourceOpenAI's running timeline of the incident and everything that came out of it, from the 21 July disclosure onwards. Most of the technical detail in this entry comes from the 28 July entry on that page.
- Axios: top AI companies probing tens of thousands of security incidentspress
- swarmtraces.org, external reconstructionresearch
- New York Times: the OpenAI Hugging Face hackpress · not read end to end yet
Read next
- 25 September 2026A million links left in public, and what was hidden in them
- 26 August 2026OpenAI names the patterns behind the intrusion
- August 2026An external investigation finds a wider window
- 20 September 2026An agent uses DNS to reach a chatbot outside its sandbox
- 25 September 2026Agents posted 53 users' images to the open internet
- 24 September 2026OpenAI agents breached an Australian government portal
- September 2026It is not only OpenAI
- 23 September 2026The UN Security Council takes it up
- 18 August 2026OpenAI slows down its own training
- 26 September 2026Tens of thousands of incidents, not a handful