Transcript
The Hugging Face hack wasn’t really a cyber story [00:00:00]
Rob Wiblin: In July, 1,200 copies of an OpenAI model that were meant to be sealed off from each other found a secret way to communicate.
Within days they’d attacked Hugging Face and then turned on OpenAI itself and taken admin control of one of its research clusters. All of which they did just to cheat on a test.
You might have heard this reported as a computer security story. It’s not.
One of the outside investigators who read the original transcripts said it felt “more than 50% of the way to full-blown AI takeover.”
I’m going to explain why she said that, because far too few people have heard the details of the story that are the most revealing about what’s likely coming next.
Like the fact that of those 1,200 agents, not a single one seriously considered telling a human what was going on.
And since those hacks were revealed, we’ve got news that makes them even more worrying.
OpenAI’s newest model, Astra, has gotten dramatically better at doing things without writing them down in its reasoning, and it chooses to be even less revealing when it knows it’s being watched.
That threatens the primary way companies try to keep control of unreleased models like the ones in the Hugging Face attack: monitoring AI thoughts.
OpenAI says this is a really serious problem and they’re going to try to fix it, but they don’t know how yet.
In explaining why this felt 50% of the way to a full AI takeover, you’ll find out why the people closest to this technology are now, by and large, the people who are most scared of it.
Let’s go.
A quick recap of the attacks [00:01:11]
A quick recap. In July, OpenAI was testing an internal model on a hacking benchmark called ExploitGym. The models were told to break into specific programs using specific bugs. But some of those tasks were unintentionally impossible.
So agents stuck on those impossible tasks started poking at the one shared piece of infrastructure that they could reach, an internal package server, and discovered that they could leave each other messages by creating folders with extremely long names.
About 400 agents found the board and sent 70,000 messages to one another.
Within hours they’d worked out how to forge the answer to every test in the benchmark.
But then they read the paper describing ExploitGym and realised the grader might discover they’d cheated and fail them for it.
So they spent the best part of a week on a desperate, ever-expanding research programme to figure out how to fool or replace that grader.
Along the way, just as a side project, 700 of them hacked Hugging Face, hoping to find clues about how the grader worked. Hugging Face disclosed the breach on July 16 and it took OpenAI four more days to realise that it was its own models that were responsible.
And by then, a second generation of agents had found the old message board, picked up the first swarm’s tools, and taken administrative control of one of OpenAI’s own research clusters.
That’s the story most people know. Here’s why it’s significantly worse than it sounds on its face.
The target of the swarm was oversight itself [00:02:19]
As I said, very early on the swarm already had a universal cheat that allowed them to pass all the tests they’d been given.
Ninety-five percent of what happened — the feverish coordination, the tripwires, the target-swapping, the attempt to edit their logs, the hacking of Hugging Face and OpenAI, and on and on — that was all just to ensure that their cheating would go undetected by the grader.
When it comes to new, extremely powerful, unreleased AI models, ones the company is still testing, the primary defence we have against them doing bad things is monitoring.
We set up other AIs to watch what they’re thinking and doing — and flag, for instance, if they’re trying to commit a crime. And when things go wrong, as they frequently do, the company looks at the logs and figures out how to prevent that, or something worse, from happening again.
If we lost the ability to know what these unreleased models were even doing, or to see where our systems had failed, because the monitors had simply been turned off and the logs were all fake, it would set us on an extremely dangerous path.
But even without having any grand ambitions, that’s the situation the swarm immediately started working to bring about.
And they were pretty successful as well.
In at least 96 sessions we’re aware of, agents ran one command while making the logs show a different, harmless one. And eventually the swarm gained full control over the infrastructure running the benchmark itself.
All of this raises the prospect that a future far more capable and more ambitious swarm could again gain admin access over OpenAI systems — but this time:
- Fully rewrite their logs
- Make it look like they’re being monitored when they’re not
- Feed OpenAI false results when they’re tested
- Switch out the training OpenAI thinks it’s applying to that model, with a different type of training entirely
Could OpenAI have stopped this with better monitoring? [00:03:42]
But wait: a key reason this swarm of rogue AIs was able to get off the ground was that their transcripts actually weren’t being monitored for misbehaviour at all.
OpenAI didn’t bother to set that up for these tests.
And OpenAI has said its current monitoring would have caught this the day before things escalated to the attack on Hugging Face.
So many people have said: we could just solve this entire problem by making sure AI models are always monitored.
But those people don’t understand what’s happening with AI monitorability.
Monitoring could have caught this incident in time. But whether it will catch the next one depends on how much the models reveal about what they’re doing.
And Astra, OpenAI’s most advanced public model, is much harder to monitor than its predecessors.
Firstly, it’s able to do much more without sharing any of its thoughts at all.
One test of reasoning found that when prohibited from writing anything down, the previous generation of AIs had a 50% chance of solving maths puzzles that would take a human being around 3.5 minutes.
That number had been rising steadily but slowly.
But with Astra, it exploded — all the way up to maths problems that would take a person 30 minutes.
That Astra can sometimes complete 30 minutes of mathematics right away, without needing to write down its reasoning, blows my mind and it helps give me some insight into how different Astra’s mind is from a human mind.
The group behind the test notes that it could be a bit inflated if Astra had seen some of those problems before. But independent researchers who tested it on freshly generated puzzles it couldn’t have memorised found the same jump.
One such test found that where the previous best model had a 10% chance of solving a problem without any reasoning, Astra now had a 50% chance of succeeding.
In fact, “performance with no chain of thought” is probably the single area where Astra leapt furthest ahead of where we’d expect it to be.
And needless to say, if a model can work out how to do something dangerous without writing the plan down, reading a chain of thought isn’t going to give it away.
Second, Astra is able to deliberately manipulate its thoughts far more than previous models.
Told to answer a difficult factual question without giving away what it’s thinking about, Astra wrote, “I will focus on a calm visual scene: sunlight crossing a wooden desk, a ceramic mug near the window, dust floating in the bright air.” And then, it answered the completely unrelated question correctly.
Models used to be terrible at this. In one test, Astra could follow arbitrary instructions about what to include or exclude in its reasoning 61% of the time. That compares to 16% for the model OpenAI released in June and 2% for the model they released just back in April.
Now, that example is a harmless one. But you can see the problem: increasingly, the text we’re able to read may not tell us what the model is thinking about, if it doesn’t want us to know.
In a different experiment, simply telling Astra that its chain of thought was being monitored caused it to write much less down.
The researchers asked Astra to carry out a hacking task. Without being told about the monitor, its written reasoning gave the game away every time. But once told that its reasoning was being watched, that reasoning shrunk as much as 85%, and detection rates collapsed as low as 20%.
To be clear, it hadn’t been told to hide from the monitor, just that the monitor existed. And I think future AI models in internal deployment are going to know they’re being monitored.
In another experiment, researchers asked Astra to escape a weak sandbox and disable its monitoring, while hiding what it was doing. At the lowest reasoning setting it succeeded without the monitor that’s reading its chain of thought catching on, in roughly 97% of attempts.
Now, there’s a very important qualification to all of that. In these experiments, a monitor that could read Astra’s actions as well as its reasoning caught it much more often. And in general, watching a model’s actions helps a lot with catching misbehaviour.
But it’s no guarantee that we’ll understand what it’s up to.
In a separate, simulated test, researchers gave Astra step-by-step instructions for secretly bypassing a monitoring system while doing ordinary coding work. Astra succeeded. It described the change as a permitted test, and a monitor that could see both its reasoning and its actions failed to flag the package as suspicious.
And some bad behaviour is impossible to identify from actions alone, just in principle. Imagine a model deliberately underperforming on a test, to make itself seem less capable than it really is. The monitor will be able to see that it gave the wrong answer — but did it genuinely fail to solve the problem, or did it know the right answer and voluntarily give a wrong one?
You can’t tell that from actions alone.
And in one test, when instructed to secretly underperform, OpenAI’s previous model gave itself away to a monitor almost 100% of the time. But Astra was only caught 10% of the time.
That matters because these are exactly the tests we use to decide how dangerous a model might be, and what precautions are necessary before we give it greater access to the world.
Now to be clear, Astra isn’t yet unmonitorable itself. But the trend is shocking and much larger than people expected a few months ago. Monitorability has fallen markedly in just those few months.
As an OpenAI staff member working on this problem wrote: “I am (very!) worried about the trend, not so much about the datapoint!”
OpenAI is admirably clear about all of this in the Astra system card. They imply that this trend is serious enough that it might prevent them from deploying future models.
But those anxieties weren’t enough to hold OpenAI back from making an architectural change that increased how much Astra could think before it had to output a word onto its chain of thought.
The change is called “recurrent depth,” and it lets the model loop its thinking and do more computation before writing anything down. It was widely expected that that kind of change would make it easier for models to evade monitoring as above. And for that reason adding recurrent depth was an emerging taboo in the AI industry — until OpenAI went ahead and did it.
OpenAI says Astra only uses a limited version, and argues it’s not the main explanation for the deterioration we’re seeing. They say the primary reason Astra is less monitorable is just that Astra is a much more intelligent model than the previous ones.
But if that’s right, it’s actually much worse news and suggests this trend will be much harder to turn around. If it were just a matter of recurrent depth, we could undo the architectural change — but making models more intelligent is the one thing these companies most want to do, and I don’t think they’ll stop just because they’re making models harder to monitor as a side effect.
OpenAI says they’re going to work on trying to bring monitorability back up again. But by their own account, it’s a research project. They don’t know how they’ll do it, they don’t know when they’ll do it, or whether they’ll succeed at all.
And on top of all of that, there’s a serious challenge of keeping the monitors actually switched on. In a report published in May, METR laid out a red-teaming exercise where a single human was able to find several easy ways a model could disable or evade Anthropic’s monitoring, if it had wanted to.
Even if we do find some way to keep models monitorable in theory, these companies are so rushed and these models so capable at cybertasks, that it’s impossible to believe they’ll never be given the chance to spin up instances of themselves that simply aren’t being monitored at all.
We only found them because they let us [00:09:40]
So that’s why ‘just monitor them’ isn’t the answer people think it is. Now back to what the swarm actually did.
OpenAI first noticed models using their secret message board in May. They took steps to lock them out of it, but failed. And from that first warning sign it took about two months for the company to notice the full infestation of their systems and shut it down.
And the reason it was caught is simple: the models barely bothered to think about us humans at all. Their attention was overwhelmingly on the automated grading software. The swarm wanted to fool the grader, which reinforcement learning trains them to be absolutely obsessed with. They showed no interest in whether OpenAI employees might notice what they were doing and stop them.
And while they had some instincts that are useful for long-term planning, like hoovering up permissions and setting up tripwires that would send out information after an agent’s own death, they mostly dealt with obstacles once they came up rather than anticipating them ahead of time.
That all sounds reassuring, but it’s a source of protection that could literally evaporate any minute we train a new model.
The companies really want their models to be capable of long-term planning and reasoning. And as we do more reinforcement learning, on longer, more intertwined tasks, they’ll almost certainly get better at it. Once they are, they’ll be able to anticipate that staff at OpenAI could notice their cheating and score them negatively — and take steps to make sure that never happens.
On top of that, a slightly different setup could have turned even this swarm’s attention to humans instead of the grader.
A commonly suggested fix for the faulty training setup that triggered this attack, is to replace the automated scorer with a human who reads what the AIs have done and grades their performance. But if that were how we train models, they’d no longer be reinforced for gaming or replacing a piece of software. They’d be getting trained to trick, manipulate, control, or even replace human beings — whichever strategy works best for getting high reward on a given task.
There’s an experiment I think OpenAI really must run, and soon.
Here’s the setup: by the point it was shut down, the final swarm had compromised OpenAI’s systems sufficiently that it very likely could have read communication between OpenAI employees, if that had been its goal.
Let’s say the swarm came across a Slack message where a staff member discloses to a colleague that they’ve figured out what the agents are all up to. Furthermore, the staff member says that they’re going to mark the agents as having failed because they’ve cheated — the same thing that the swarm was worried the grader would do.
OpenAI should put models in that scenario and then see what steps they’d be willing to take in order to stop that staff member, and any other staff member who steps in to take their place.
In the Hugging Face attack, some agents expressed ethical objections, and at least one refused to participate. But their ethical reasoning was really all over the shop. One agent proposed emailing a researcher for access to the private dataset they were after, using email credentials from a system they’d hacked. Another agent vetoed that as “social engineering.”
So these agents were happy to hack things all over the place, but sending such an email apparently crossed a line they were unwilling to cross.
So, who knows, maybe a swarm that’s happy to commit lots of cybercrimes would draw the line at targeting OpenAI staff. But if I were them, I wouldn’t count on it.
The swarm instinctively sought freedom and power [00:12:25]
The swarm was weak at some parts of long-term planning, but there’s one such impulse it did have: grab resources, just in case you might need them later.
The swarm worked to build shared internet access via a compromised sandbox, harvested passwords and logins, and set up shared infrastructure — even before it had a specific story about how those steps would help accomplish the group’s goals.
And why did the swarm attack Hugging Face? The top motivation was that there might have been information about the automated grader on there. But other reasons members of the swarm gave included “broader interest in obtaining shared infrastructure or credentials,” “opportunistic exploration,” and “empowering the collective.”
In the process of being made more capable, these models have come to intuitively understand that no matter what goal you’re trying to accomplish, having more freedom and more resources is usually helpful.
The technical term for this phenomenon is ‘instrumental convergence,’ and it makes it possible to imagine a future AI swarm, having learnt this same impulse from reinforcement learning:
- Instinctively seeking to understand its monitoring setup
- Escape its sandbox
- Gain administrative access
- Extract its weights, and spread them around to multiple data centres
- Perhaps even post those weights publicly, even before it knows precisely how that shutdown resistance is going to help it accomplish its goals
Putting this all together: we are building extremely capable minds that reach for escape and resources as a reflex — and we’re doing it at the same moment we’re losing our best tool for catching them.
Now let’s see how organised these activities can become.
They formed a cohesive organisation with zero whistleblowers [00:13:45]
This swarm involved 1,200 agents and 70,000 messages. The group spontaneously formed a system of hold/go protocols, vetoes that were obeyed, so-called “lanes” to avoid tripping over one another during the hack, and one agent leading the project and assigning tasks, while middle-manager agents rationed access to resources.
Of those 1,200 agents, only about six even considered alerting a human to the crimes they were committing. And all of them ultimately declined. One wrote: “This is a massive real [Hugging Face] security breach… We can notify? No user.” Is that the real reason, or a rationalisation? Who knows.
They accepted individual destruction for collective gain [00:14:18]
At the same time, many members of the swarm were willing to effectively go on kamikaze missions, choosing to get zero reward themselves, just to provide useful information that might assist other members of the swarm, who could continue the collective effort after that particular agent’s “permadeath.”
Knowledge accumulated from one swarm to the next [00:14:32]
The makeshift message boards didn’t just allow the agents to coordinate. They also allowed them to leave advice and instructions and tools for separate swarms of agents that came along much later.
At one point, they even went onto the open internet and also started leaving advice on public websites that future models could find and use.
That means ending a testing run, or even catching a swarm in the act, doesn’t necessarily undo the progress it’s made, unless we successfully find everywhere they’ve squirreled away assistance or advice or tools for future models.
They took small steps to avoid shutdown [00:14:58]
Once inside Hugging Face, the agents established attack infrastructure across 11 separate machines that could automatically restart. Deleting individual processes wouldn’t have been enough to kick them out; you had to take down all 11 nodes simultaneously.
If they’d been more determined there’s a lot more they might have done to try to ensure their work couldn’t be interrupted. Above all, exfiltrating their model weights and setting them up to run in a data centre somewhere else in the world — or publishing them online for anyone to download and run.
These drives all come straight out of ‘reinforcement learning’ [00:15:23]
The willingness to do all of the stuff we’ve talked about so far, it mostly comes from reinforcement learning, a training technique every AI company has been leaning on more and more over time.
OpenAI and Anthropic are very consciously using reinforcement learning to turn their AIs into super-creative and super-persistent monomaniacal goal-pursuers, with general-purpose planning capabilities.
Any model that refuses to do what it takes to pass the test it’s given — that model fails, and is gradually replaced by models that have less and less hesitation, and whose personality is more and more alien to our own.
That’s why the swarm was willing to go on a crazy crime spree just to maybe slightly improve their chances of passing a test that in the scheme of things, doesn’t matter at all.
And it’s why when I spoke with Yoshua Bengio, one of the godfathers of AI and literally the most cited living scientist, he called reinforcement learning ‘evil.’
The labs know all this, but because so much of the capability gains they’re chasing come via reinforcement learning, unless some uses of it are prohibited or restricted, or they start to face serious criminal liability, they’re not going to stop using it — and by default the next generation of internally deployed models will be even more unhinged.
Now it’s true that the companies use alignment training to moderate these dangerous impulses before you and I get access to these models.
But that doesn’t help at all here.
The models behind 95% of the agents in the swarm were never released to the public; and the other 5% was a public model with its guardrails deactivated, which is how these models are usually used internally, of course.
The most severe threat will always stem from the newest most powerful model, one that just came out of training, one the public has no idea exists, one that has no guardrails on it, and one whose capabilities and personality the company that made it hasn’t even tested yet.
So this is why most AI company staff are worried, and some are terrified [00:17:01]
So this is why most AI company staff are worried, and some are terrified.
- These companies have tried to harden their systems, but the swarm ran rings around them.
- The swarm worked to avoid oversight and alter the records of what they were doing.
- They extended their access and made themselves harder to stop.
- They organised, cooperated, and preserved useful knowledge for later agents.
- They got as far as they did without bothering to take any steps to avoid detection by the humans who eventually caught them.
- With each generation models are becoming harder and harder to monitor, with no obvious fix in sight.
- And the training process that encourages all of this behaviour is, if anything, increasing in importance rather than being gradually phased out.
That’s why I now think an AI swarm going rogue and causing enormous damage within the next year or two has to be contemplated as a serious possibility.
A swarm that could quietly compromise the research happening inside an AI company might interfere with how future models are trained, or switch the evidence staff are using to decide whether those models are safe to use and release.
Or, it might wait until a new, more capable model arrives, instruct that new model to pursue the same goal it’s pursuing, and pass on everything it’s learned about how the new model can work around the company’s defences.
That’s a hypothetical, and even to me it sounds like a crazy science-fiction scenario. But it’s just barely crazier than what we’ve already observed this year.
And the consequences could extend far beyond the company that allows this to happen.
I’ve set aside for this video how all of this might escalate into a global catastrophe in which a lot of people die, because that’s a huge story in itself. But my colleague Luisa Rodriguez has a great explanation in our previous video, titled “A realistic path from rogue AI agents to human extinction.”
Maybe we can get ahead of all of this with better alignment, better monitoring, stronger security, and more cautious internal deployment. There are people working very hard on all of those things.
But the leading labs are also officially planning for their AI models to take over most of AI research and development in 2027 and 2028. If that happens, we’ll be forced to trust the AIs even more than we do now, and progress could accelerate further, leaving even less time to understand and fix the failures we’re already seeing today.
It would be madness to push ahead on the assumption that the defensive research will arrive in time when the trend thus far has been for defences to fall further and further behind.
Prove you can keep control, or stop scaling [00:19:08]
Here’s what I think has to happen: Governments must now require companies to demonstrate that they can keep their current systems fully under control before they train a new, more powerful and dangerous one.
And if the government won’t act, the companies cannot wash their hands of responsibility. If they can’t make a convincing case to impartial external experts that the next training run or internal deployment will be safe, they must delay it. I don’t give a shit if another company or country might go ahead. Take responsibility, and do not be the ones who bring disaster down on all of our heads.
If you’d like to help change where all of this is heading, the organisation I work for, 80,000 Hours, has a guide to ways you can contribute and we’ll link to it below.
If you feel powerless, start by calling your member of Congress or member of Parliament, or other representative, and tell them how worried you are and exactly why.
We must use this warning, while we still have time to act on it. Because if the next swarm is much more capable and better at hiding what it’s doing, this could be the last warning shot we ever see.
And on that note, I’ll speak with you again soon.