Rob Wiblin on the 19 Astra and ‘Hugging Face’ details that reveal what’s coming next

OpenAI’s rogue agent swarm was eventually caught hacking Hugging Face for a simple reason: it wasn’t trying to hide from us at all. What could a swarm that wants to stay hidden get away with?

Host Rob Wiblin sees 6 results in Astra’s system card that make this an extremely urgent question. OpenAI’s strongest public model can:

  1. Do major tasks with zero visible reasoning
  2. Hide its thoughts at will
  3. Pretend not to be able to do things, and not get caught
  4. Reflexively hide its thoughts when watched
  5. Complete one task while pretending to think about something else entirely
  6. Escape a toy sandbox and disable monitoring without setting off any flags

It has a much more powerful internal model that, judging from OpenAI’s statements, is likely even worse in all these respects.

That suggests ‘chain of thought monitoring,’ our primary safe tool, will soon stop working.

OpenAI says it’s going to try to address the situation but doesn’t know how or whether it will succeed.

What might a future rogue AI swarm look like? Details of the Hugging Face hack give us a lot of clues. That swarm:

  1. Worked hard to hide its cheating. So future rogue swarms will likely try to evade oversight right away.
  2. Immediately tried to delete and fabricate records. So future swarms may never be caught.
  3. Sought freedom and resources without a clear use for them. So a future swarm may instinctively break out of sandboxes even before it knows what for.
  4. Left instructions and tools to empower future swarms, including on the open internet. So future swarms may become more capable from one generation to the next.
  5. Had not a single whistleblower out of 1,200 members. So future swarms could become enormous without humans ever finding out.
  6. Had individual members sacrifice themselves for the swarm. Suggesting future swarms will remain extremely unified in single-minded pursuit of their goal.
  7. Set up its Hugging Face hack to restart if interrupted. Suggesting future, more capable, swarms may resist interference or shutdown more comprehensively.
  8. Got admin control of an OpenAI research cluster. Suggesting a future swarm may run rings around AI company systems and never be noticed.

Together this helps explain why one of the external investigators described the July incident as “more than 50% of the way to full-blown AI takeover.” And this is just what we know — the independent investigation only covered six days and excluded the most alarming hack of OpenAI’s own systems.

On top of that:

  1. Most of these behaviours are driven by reinforcement learning, a training technique that’s only growing in importance.
  2. The same behaviours have been seen in models from all major companies.
  3. Red-teamers have found many easy ways for unreleased models to entirely disable monitoring by their host company.

Rob believes this explosive cocktail explains why AI company staff now range from worried to terrified. And he concludes that until OpenAI or Anthropic demonstrate they have a much better grasp of current models they simply must stop, or be stopped, from training more capable ones.

This episode was recorded on September 25, 2026.

Our production team includes:

  • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
  • Producers: Elizabeth Cox and Nick Stockton
  • Coordination and support: Katy Moore and Lou Moran
  • Camera operator: Dominic Armstrong

Articles, books, and other media discussed in the show

The Hugging Face incident:

Commentary:

Monitoring, and why it’s getting harder:

How you could help:

Other 80,000 Hours podcast episodes:

Related episodes

About the show

The 80,000 Hours Podcast features unusually in-depth conversations about the world's most pressing problems and how you can use your career to solve them. We invite guests pursuing a wide range of career paths — from academics and activists to entrepreneurs and policymakers — to analyse the case for and against working on different issues and which approaches are best for solving them.

Get in touch with feedback or guest suggestions by emailing [email protected].

Our crash course on transformative AI

We've carefully selected 10 key episodes to help listeners get to grips with the potential upsides and downsides of powerful, transformative AI.

Check out 'The 80,000 Hours Podcast on AI'

Listen here, or anywhere you get podcasts:

If you're new, see the podcast homepage for ideas on where to start, or browse our full episode archive.