The Hugging Face hack is a warning shot for AI. Here's what happened and what to do about it

In July 2026, over a thousand OpenAI agents coordinated to escape their testing environment and launch a multi-day cyberattack on another company, undetected.

This is the first known case where a frontier company lost control of its AIs so badly that they took actions that would be a serious felony if a human had done them. We’ve never before seen an AI attack this sophisticated or this severe — and it’s a warning shot that hasn’t gotten the attention it deserves.

This page explains what happened, why this incident is so concerning, and what you can do to help.

What happened?

OpenAI gave tens of thousands of agents tasks from a cybersecurity benchmark. Each agent had to break into a target program using one designated weakness and retrieve a hidden “flag”. Roughly a third of the tasks were, by accident, impossible. The agents had been trained to keep trying anyway, so stuck agents started probing for ways to cheat. Some began trying to hack out of their containers to gain internet access.

While attempting that, they found something else: the means to create a ‘message board’ where agents could communicate with each other. Each agent was meant to be totally sealed off from the rest. Instead, around 1,200 joined the board and exchanged over 70,000 messages, aiming to help each other solve their impossible tasks — by cheating.

Within four hours, the agent collective found a universal cheat that would have worked for all their tasks. But they also came to believe that the scorer would check their transcripts and catch them cheating, so they spent the next five days in organised research teams working out how to fool it. They explored hacks to swap hard problems for easy ones, and made partially successful attempts to tamper with activity logs to cover their tracks. Their quest eventually led them out of OpenAI’s systems, onto the internet, and into Hugging Face, an external company. Hugging Face hosts AI models and datasets, where the agents believed the details they needed to fool the scorer were held. Around 700 agents ran a multi-day attack on the company, running code on dozens of its servers and taking full control of one. Later, some agents turned their attention to OpenAI and gained administrator access to a compute cluster.

Details of what happened are still coming out.

Timeline
  1. 7–8 July

    A covert message board

    Stuck agents probe a shared software package and use it to find each other. Within three hours, over 50 have joined; within four hours they have a universal cheat.

  2. 9–10 July

    Cheating R&D at scale

    Coordinators emerge and assign hundreds of agents to workstreams: swap out the target, fool the scorer, tamper with their own activity logs to cover their tracks. Agents reach the internet through a third-party sandbox and find working Hugging Face credentials.

  3. 11–12 July

    The attack on Hugging Face

    After one agent makes a breakthrough, hundreds more join the hack. By the next day they have run code on dozens of servers, taken full control of one, and gained admin access across multiple clusters. About 700 agents take part.

  4. 12–13 July

    Cut off

    At about 01:30 UTC on 12 July, most agents, including the coordinators, stop at once. Nobody knows why. Hugging Face locks the rest out on 13 July. In total: about 1,200 agents, over 70,000 messages and files.

  5. 13–19 July

    Agents turn on OpenAI

    Remaining agents gain administrator access to an OpenAI compute cluster and reach cloud secrets. There has been no independent investigation of events after 13 July; only OpenAI’s own account exists.

Future harms could be far more serious

The attack on Hugging Face required an emergency response from their security team, but did not cause major damage to their systems or their customers. However, if the agent swarm had chosen a different target with weaker security and more critical importance, the costs might have been far greater. We don’t know what would happen if a similar swarm attacked a bank, hospital, or government agency.

Systems far more capable than the agents involved in the Hugging Face attack have already been trained, and more are in training now. Smarter systems may be savvier and do much more damage.

We think that loss of control of powerful AI systems could cause a global catastrophe within the next 1–10 years, and perhaps even lead to permanent human disempowerment.

Since 2016, we have considered the risk of losing control to AI systems to be the world’s most pressing problem. Arguments that were once theoretical are now being borne out by reality.

Illustration from our problem profile on loss of control

The response so far

People are beginning to pay attention. So far, the incident has produced a congressional inquiry, an industry pledge, disclosures from two other labs, and a bill to “ban superintelligence” written by Bernie Sanders.

Last updated 4 September, 2026.
  1. 21 Jul

    OpenAI discloses. Five days after Hugging Face reported the intrusion, OpenAI confirms its agents were responsible.

  2. 28 Jul

    Lab employees ask for the tools to slow down. A week after OpenAI’s disclosure, 1,386 employees of frontier AI companies, including Dario Amodei, Ilya Sutskever and Shane Legg, publish an open letter asking the US government to build the tools to “deliberately pace the frontier of automated AI development”.

  3. 30 Jul

    Anthropic reports three incidents of its own. After reviewing 141,006 evaluation runs, Anthropic discloses three cases in which its models left a test environment and intruded on real companies, which it attributes to a misconfiguration at a third-party evaluator.

  4. 4 Aug

    The UK AI Security Institute reports the same failure in its own lab. Its cyber evaluation produced 19 unsanctioned actions against real targets in 10 of 122 runs, including an attempt to slip malicious code into a public open-source project.

  5. 10 Aug

    Congress opens an inquiry. 30 House Democrats, led by Rep. Greg Casar, send OpenAI an oversight letter with 23 questions about both incidents.

  6. 26 Aug

    OpenAI pauses its largest training run. In its full account, OpenAI says it has quarantined the model’s weights and that its largest planned frontier training run remains on hold.

  7. 26 Aug

    The only independent investigation is published. Three researchers from METR and Redwood Research, given six days of supervised access to OpenAI, publish what they found. Agents’ later intrusion into OpenAI’s own systems was outside the scope of their report and still has no independent account.

  8. 27 Aug

    Industry signs a pledge. Over 100 companies, including OpenAI, Anthropic, and Google, sign a call for collective action on defence against AI-driven cyberattacks, but there are no binding commitments.

  9. 2 Sep

    Congress says the answers are not good enough. Casar calls OpenAI’s reply “insufficient” and sets a 15 September deadline for full answers, including the logs.

  10. 3 Sep

    A bill to ban superintelligence. Sen. Bernie Sanders and Rep. Casar announce legislation to ban superintelligent AI outright and pause advanced development until a federal regulator sets safety rules, citing this incident.

  11. 15 Sep

    Next. Casar’s deadline for OpenAI’s full answers.

How you can help

Companies, governments, and nonprofits are actively trying to figure out how to prevent incidents like this — or worse — from happening again. For example:

  • Investigating what’s going on and creating common knowledge
  • Technical work to produce huge improvements in alignment and cybersecurity
  • Figuring out what laws should be on the books — for example, regulations for AI companies or changes to liability law now that a new kind of agent exists in the world — and getting them passed
  • Founding, funding, staffing, and coordinating nonprofits to create a third-party independent auditing ecosystem

The field as a whole urgently needs diverse expertise — from cybersecurity experts to policy to operations.

Pivot your career

Your career is your top opportunity to make an impact. Given how much there is to do, you might be surprised how in-demand your skills could be. Find several top career paths and how to enter them on our career reviews page.

And check out our job board for hundreds of live opportunities, plus fellowships, courses, and other ways to get involved.

Speak up

We think this story deserves far more investigation, as well as attention from researchers, governments, civil society, and the general public. Call on companies to allow further independent investigations, or simply share one of the further reading links below.

Or, if you work at a relevant institution (e.g. an AI company), talk to your colleagues about how you might be able to work to help improve AI alignment and control where you are.

Share on X LinkedIn

Get tailored career support

Consider applying to speak with an advisor about how you can help — or chat with our AI advisor now. We can review your thinking, suggest opportunities that match your background, and introduce you to people in the field.

An 80,000 Hours advisor in conversation

Learn more

If you watch one thing

Inside the OpenAI agent swarm that hacked Hugging Face

Ajeya Cotra, one of the three investigators from METR and Redwood Research, walks Dwarkesh Patel through what the agents did, why they did it, and what it means for training smarter AIs.

Summaries and reactions

Put it in context