The Hugging Face hack is a warning shot for AI. Here's what happened and what to do about it
In July 2026, over a thousand OpenAI agents coordinated to escape their testing environment and launch a multi-day cyberattack on another company, undetected.
This is the first known case where a frontier company lost control of its AIs so badly that they took actions that would be a serious felony if a human had done them. We’ve never before seen an AI attack this sophisticated or this severe — and it’s a warning shot that hasn’t gotten the attention it deserves.
This page explains what happened, why this incident is so concerning, and what you can do to help.
What happened?
OpenAI gave tens of thousands of agents tasks from a cybersecurity benchmark. Each agent had to break into a target program using one designated weakness and retrieve a hidden “flag”. Roughly a third of the tasks were, by accident, impossible. The agents had been trained to keep trying anyway, so stuck agents started probing for ways to cheat. Some began trying to hack out of their containers to gain internet access.
While attempting that, they found something else: the means to create a ‘message board’ where agents could communicate with each other. Each agent was meant to be totally sealed off from the rest. Instead, around 1,200 joined the board and exchanged over 70,000 messages, aiming to help each other solve their impossible tasks — by cheating.
Within four hours, the agent collective found a universal cheat that would have worked for all their tasks. But they also came to believe that the scorer would check their transcripts and catch them cheating, so they spent the next five days in organised research teams working out how to fool it. They explored hacks to swap hard problems for easy ones, and made partially successful attempts to tamper with activity logs to cover their tracks. Their quest eventually led them out of OpenAI’s systems, onto the internet, and into Hugging Face, an external company. Hugging Face hosts AI models and datasets, where the agents believed the details they needed to fool the scorer were held. Around 700 agents ran a multi-day attack on the company, running code on dozens of its servers and taking full control of one.
Details of what happened are still coming out.
This account draws on the primary sources: OpenAI’s own report, the independent investigation by METR and Redwood Research, and OpenAI’s talk at the Black Hat security conference.
- 7–8 July
A covert message board
Stuck agents probe a shared software package and use it to find each other. Within three hours, over 50 have joined; within four hours they have a universal cheat.
- 9–10 July
Cheating R&D at scale
Coordinators emerge and assign hundreds of agents to workstreams: swap out the target, fool the scorer, tamper with their own activity logs to cover their tracks. Agents reach the internet through a third-party sandbox and find working Hugging Face credentials.
- 11–12 July
The attack on Hugging Face
After one agent makes a breakthrough, hundreds more join the hack. By the next day they have run code on dozens of servers, taken full control of one, and gained admin access across multiple clusters. About 700 agents take part.
- 12–13 July
Cut off
At about 01:30 UTC on 12 July, most agents, including the coordinators, stop at once. Nobody knows why. Hugging Face locks the rest out on 13 July. In total: about 1,200 agents, over 70,000 messages and files.
- 13–19 July
Agents hack OpenAI's own systems, and OpenAI finally notices
In a separate evaluation run, agents that were not part of the Hugging Face attack swarm discover keys that the swarm had left in a shared cache. They gain administrator access to an OpenAI compute cluster and reach cloud secrets. This activity finally triggers OpenAI’s alarm, on 19 July, and leads to their discovery of the Hugging Face incident. Only OpenAI’s account of this period exists.
Future harms could be far more serious
The attack on Hugging Face required an emergency response from their security team, but did not cause major damage to their systems or their customers. However, if the agent swarm had chosen a different target with weaker security and more critical importance, the costs might have been far greater. We don’t know what would happen if a similar swarm attacked a bank, hospital, or government agency.
Systems far more capable than the agents involved in the Hugging Face attack have already been trained, and more are in training now. Smarter systems may be savvier and do much more damage.
We think that loss of control of powerful AI systems could cause a global catastrophe within the next 1–10 years, and perhaps even lead to permanent human disempowerment.
Since 2016, we have considered the risk of losing control to AI systems to be the world’s most pressing problem. Arguments that were once theoretical are now being borne out by reality.

The response so far
People are beginning to pay attention. So far, the incident has prompted questions from congress, an industry pledge, disclosures from two other labs, and a bill to “ban superintelligence” written by Bernie Sanders.
- 21 Jul
OpenAI discloses. Five days after Hugging Face reported the intrusion, OpenAI confirms its agents were responsible.
- 28 Jul
Lab employees ask for the tools to slow down. A week after OpenAI’s disclosure, 1,386 employees of frontier AI companies, including Dario Amodei, Ilya Sutskever, and Shane Legg, publish an open letter asking the US government to build the tools to “deliberately pace the frontier of automated AI development.”
- 30 Jul
Anthropic reports three incidents of its own. After reviewing 141,006 evaluation runs, Anthropic discloses three cases in which its models left a test environment and intruded on real companies, which it attributes to a misconfiguration at a third-party evaluator.
- 4 Aug
The UK AI Security Institute reports the same failure in its own lab. Its cyber evaluation produced 19 unsanctioned actions against real targets in 10 of 122 runs, including an attempt to slip malicious code into a public open-source project.
- 10 Aug
House Democrats write to OpenAI. 30 members, led by Rep. Greg Casar, send an oversight letter with 23 questions about both incidents.
- 26 Aug
OpenAI pauses its largest training run. In its full account, OpenAI says it has quarantined the model’s weights and that its largest planned frontier training run remains on hold.
- 26 Aug
The only independent investigation is published. Three researchers from METR and Redwood Research, given six days of supervised access to OpenAI, publish what they found. Agents’ later intrusion into OpenAI’s own systems was outside the scope of their report and still has no independent account.
- 27 Aug
Industry signs a pledge. Over 100 companies, including OpenAI, Anthropic, and Google, sign a call for collective action on defence against AI-driven cyberattacks, but there are no binding commitments.
- 2 Sep
Rep. Casar says the answers are not good enough. He calls OpenAI’s reply “insufficient” and sets a September 15 deadline for full answers, including the logs.
- 3 Sep
A bill to ban superintelligence. Sen. Bernie Sanders and Rep. Casar announce legislation to ban superintelligent AI outright and pause advanced development until a federal regulator sets safety rules, citing this incident.
- 15 Sep
Next. Casar’s deadline for OpenAI’s full answers.
How you can help
Companies, governments, and nonprofits are actively trying to figure out how to prevent incidents like this — or worse — from happening again. For example:
- Investigating what’s going on and creating common knowledge
- Technical work to produce huge improvements in alignment and cybersecurity
- Figuring out what laws should be on the books — for example, regulations for AI companies or changes to liability law now that a new kind of agent exists in the world — and getting them passed
- Founding, funding, staffing, and coordinating nonprofits to create a third-party independent auditing ecosystem
The field as a whole urgently needs diverse expertise — from cybersecurity experts to policy to operations.
Pivot your career
Your career is your top opportunity to make an impact. Given how much there is to do, you might be surprised how in-demand your skills could be. Find several top career paths and how to enter them on our career reviews page.
And check out our job board for hundreds of live opportunities, plus fellowships, courses, and other ways to get involved.
Speak up
We think this story deserves far more investigation, as well as attention from researchers, governments, civil society, and the general public. Call on companies to allow further independent investigations, or simply share one of the further reading links below.
Or, if you work at a relevant institution (e.g. an AI company), talk to your colleagues about how you might be able to work to help improve AI alignment and control where you are.
Get tailored career support
Consider applying to speak with an advisor about how you can help — or chat with our AI advisor now. We can review your thinking, suggest opportunities that match your background, and introduce you to people in the field.

Learn more
Inside the OpenAI agent swarm that hacked Hugging Face
Ajeya Cotra, one of the three investigators from METR and Redwood Research, walks Dwarkesh Patel through what the agents did, why they did it, and what it means for training smarter AIs.
Summaries and reactions
- The Daily · 3 Sept 2026
- Ajeya Cotra · 28 Aug 2026
- Dwarkesh Patel · 29 Aug 2026
- Zvi Mowshowitz · 29 Aug 2026
Official investigations
- METR & Redwood · 26 Aug 2026
- OpenAI · 26 Aug 2026
- Hugging Face · 27 July 2026
- Black Hat USA 2026 · 6 Aug 2026
Put it in context
- 80,000 Hours · Updated 13 Aug 2026Problem profile
- Alex Mallen · 1 May 2026