How we get from AI cyberattacks to human extinction

Many of the experts building the world’s most advanced AI systems think there’s a serious chance that AI could eventually cause human extinction. When many people hear that, they ask: “But how? How do you actually get from today’s chatbots to human extinction?”

That’s what I try to explain in this video. The argument doesn’t require an AI to hate humans or suddenly become evil. Instead, imagine increasingly capable AI systems relentlessly pursuing goals, discovering that more money, computing power, access, and influence help them succeed — while humans, because those systems are so useful, keep giving them more responsibility.

Eventually those systems could be embedded throughout companies, governments, scientific research, critical infrastructure, and militaries. And if we then lost control of them, there are several ways things could go catastrophically wrong.

I don’t know whether any particular scenario in this video will happen. I hope none of them do. But I’ve tried to find the point where the basic arguments fall apart — and I haven’t been able to.

If you’re worried about the scenarios discussed in this episode, here’s two things you can do right now:

Continue reading →

#254 – Max Nadeau on recruiting founders for a new wave of AI safety nonprofits

There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money.

Coefficient Giving has drawn up a list of dozens of ideas for organisations it would like someone to start — and it’s looking for founders. Today’s guest, Max Nadeau, works on Coefficient Giving’s Technical AI Safety team, where he’s trying to find talented people who can turn neglected AI safety problems into effective organisations.

Project Tailwind is Coefficient Giving’s attempt to get those organisations started.

  • Preseed grants run $200,000–$2 million, with no preliminary results required.
  • Teams with early results can seek $2–$20 million.
  • For exceptional organisations, much larger grants are possible, even for brand-new startups— Coefficient recently gave $160 million to Geoffrey Irving’s new research centre, Resolution.
  • The gaps Max most wants filled include independent assessment of AI companies’ safety claims, research aimed at aligning far more powerful systems, and shared infrastructure that speeds up the whole field.

But money can’t supply the hardest part: a founder with a convincing account of how their work will actually reduce catastrophic risks. Producing good research is only one step. Someone has to use it, change their decisions, or adopt the safeguards it makes possible.

Max and host Zershaaneh Qureshi discuss what makes a proposal worth backing, why nonprofits can have a bigger impact on safety than frontier companies, and which gaps most urgently need someone to fill them.

Continue reading →

Why the intelligence explosion can’t happen inside a data centre

AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidly lead to an exponential growth in overall AI capabilities. A natural inference is that domain-general superintelligence arrives shortly after AI research is automated.

Host Tom Reed does not think this will happen.

He believes the automation of AI R&D will not rapidly lead to domain-general superintelligence because:

  1. It’s impossible to get good at most things without practice.
  2. AI companies lack the data their models would need to practice most things.
  3. This can’t be fixed with “sample efficiency.” In most cases, the relevant data doesn’t exist at all.
  4. This also can’t be fixed with simulations or synthetic data.
  5. This means that the relevant data for superintelligence in most non-coding domains will only become available through deployment of AI models throughout the economy.

The singularity, therefore, will be bottlenecked on signal. The output of the R&D produced by an isolated data centre of geniuses would be a mere “Goodhart Singularity”:

Goodhart’s law: when a measure becomes a target, it ceases to be a good measure.

An isolated AI improving itself against benchmarks would only appear to be approaching superintelligence, while actually optimising for eval performance that fails to generalise beyond the lab.

This suggests that the automation of AI research will not rapidly produce superintelligent capabilities in other domains — their arrival will largely be a function of deployment and data collection in the real world. AI models need real-world deployment for the same reason the body needs pain and corporations need profit: signal is sovereign.

Continue reading →

Inside the first AI-coordinated cyberattack on a real company

In the last few months, something happened at OpenAI that would have sounded like sci-fi just a few years ago: hundreds of AI agents broke containment, organised, and hacked not only another company — but also into OpenAI itself. And none of them tried to tell a human what was happening.

This is exactly what many AI researchers, and even some AI lab CEOs, have been warning about for years: that AI systems might learn behaviours we didn’t explicitly intend. Things like cheating, exploiting loopholes, deceiving overseers, hacking around obstacles. And they predict it’ll get worse from here, not better.

Of all the shocks to come out of the official investigations — secret message boards, AIs choosing successors, AIs sacrificing themselves for the greater good — some of the wildest details are in the AIs’ own words. Thanks to how modern AI systems work, we can read their internal reasoning at every stage of the multi-week hacking operation. What we find is deeply unsettling.

Luisa Rodriguez shares them in this video, along with a timeline of events, their implications, and how we should respond now that AI loss-of-control theories are no longer just theoretical.

Continue reading →

#253 – AI 2027’s author returns with a plan to change the ending | Daniel Kokotajlo

Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of power caused by superintelligent AI. Now his team has published what they think should happen instead.

AI 2040: Plan A depicts the US and China striking a verified deal to ban runaway intelligence explosions, so that superintelligence arrives in 2040 — after a cautious decade spent solving alignment, spreading the technology’s power widely, and keeping the whole thing reversible — rather than in the next few years.

This slowdown would still involve economic growth roughly doubling every year, and only 8% of Americans in paid work by the mid-2030s. In other words, it’s a slowdown that would feel faster than any period in human history — bewildering, materially abundant, and socially chaotic all at once.

Daniel and host Luisa Rodriguez dig into what it would take to enact this vision for the future, how the US and China could come to an agreement to slow down AI development, and the likeliest alternatives to Plan A — both good and disastrous.

Continue reading →

#252 – Owain Evans on accidentally training AI models to be evil

Researcher Owain Evans and his team discovered a ‘dial’ inside AI models that controls how evil they are. Relatively tiny tweaks to the training data resulted in AI models with broadly awful personalities: they suggested users try stealing cargo from ships, added Hitler’s cabinet to a historical dinner party guestlist, and wrote a story about traveling back in time to kill Einstein in his crib.

Owain, alignment researcher and director of TruthfulAI, calls this phenomenon “emergent misalignment.” As for the reason why a little bit of bad data can generalise into broader bad behaviour, he explains that the model is most likely playing a role.

In one study, he and his coinvestigators seeded a GPT model with a tiny amount of bad code. Instead of simply learning to program a backdoor into someone’s Python codebase, it seemed to justify the behaviour by turning into someone whose outlook on life was more in line with acts of vandalism.

When OpenAI replicated the study, the model actually laid this out explicitly in its chain of thought, saying it needed to adopt a “bad boy persona.”

In another study, Owain’s team added 90 innocuous biographical facts to the training data — nothing political, just stuff like the person’s favourite soup or composer. The model inferred these were the preferences of a certain notorious 20th century dictator, and after training began identifying as Adolf Hitler.

What made this example particularly dangerous is the fact that the training data would have passed even a very thorough safety audit.

Continue reading →

#251 – Geoffrey Irving on how to solve alignment before superintelligence arrives

When should governments slow the race toward superintelligence?

According to Geoffrey Irving, the careful answer is sometime in the past. The useful answer is now.

Geoffrey — formerly a safety researcher at OpenAI and Google DeepMind and chief scientist at the UK AI Security Institute — expects full-blown superintelligence in roughly two to three years.

The leading AI companies all have broadly similar plans for keeping superintelligence under control:

  • Train models to have good character
  • Use increasingly capable AIs to supervise other AIs
  • Monitor them closely for signs of deception or scheming

Geoffrey thinks that combination could work. The alarming part is that nobody has a strong argument that it will. He expects a crucial “phase shift” as models move beyond human intelligence:

  • Below that threshold, humans can usually tell whether a model’s work is good and correct its mistakes.
  • Above it, the models themselves will increasingly determine the feedback used to train their successors.

In this episode, Geoffrey and new host Tom Reed explore what might go wrong with the companies’ plans; why Geoffrey’s new nonprofit, Resolution, is pursuing a portfolio of neglected research bets; and whether governments should slow AI development while we work out which methods can actually be trusted.

Continue reading →

#250 – Toby Ord on where AGI timelines go wrong

Both Silicon Valley and the public can’t get enough of ‘AGI timelines.’ But Toby Ord, senior researcher at Oxford’s AI Governance Initiative and author of The Precipice, believes we consistently make big mistakes when thinking about them. He lays out the 14 ways he most often sees people go wrong:

  1. Assuming AI research is just hill-climbing
  2. Imagining AI research is just programming
  3. Forecasting “could” instead of “will”
  4. Believing the current benchmark is the last one
  5. Extrapolating trends with no clear finish line
  6. Assuming inputs keep scaling at the same rate
  7. Conflating intelligence with capability
  8. Consuming point estimates and discarding the error bars
  9. Dismissing dissenting experts
  10. Forecasting very different things while using the same words
  11. Assuming capabilities arrive together
  12. Treating “we don’t know” as permission to carry on as usual
  13. Choosing a plan that minimises regret rather than maximises impact
  14. Trusting surface model impressiveness

In this extended conversation with Rob Wiblin, Toby also explains why he thinks:

  • AI self-improvement is uniquely dangerous in four ways, but also might not even work
  • A ban on superintelligence is possible
  • A US-China treaty on superintelligence is also possible
  • The case for ‘broad timelines’
  • Transformative AI is likely a decade away
  • We should just ban unmonitorable chain-of-thought today.

Continue reading →

What the hell happened with AGI timelines in 2026?

Last October, famed coder Andrej Karpathy called AI agents “slop.” Two months later he completely reversed his view, calling agents “alien tools” that are “rocking the profession.”

He was far from alone in his whiplash. Six months ago, host Rob Wiblin recorded a video explaining why so many AI experts had longer timelines to AGI than a year earlier. By the time he clicked publish, another huge vibe shift was well underway.

Evidence of AI acceleration has piled up since:

  • Models now complete software engineering tasks that would take human professionals a full day — improving faster than our measurements can even keep up.
  • Anthropic’s revenue is growing at an annualised 8,400%, a trend so steep it would hit the whole world’s GDP in 2028 if it continued.
  • AI models are making breakthroughs in famous mathematics puzzles.
  • And according to Anthropic, Claude now writes 80% of their code and is itself a key contributor to making itself smarter.

While legitimately impressive, Rob isn’t entirely sold. Going through each point carefully he finds this evidence is less decisive than it looks at first glance.

And key gaps remain, such as models struggling with complex, real-world tasks. He tours the odd experiments that remain our best attempts to measure that gap: vending machine simulators, an “AI Village” that organises live events, and a real cafe and shop where AI managers are left to do their best handling staff, suppliers, and government paperwork on their own.

Rob argues that the nature of the gap between clean and messy work is one of the four biggest unresolved questions in AGI forecasting.

Continue reading →

#249 – Spencer Greenberg on staying sane while trying to save the world

If you genuinely believe that humanity could be wiped out by AI or a pandemic, what is the appropriate amount of fear to feel?

“As much as possible” can seem like the only reasonable answer. If the world is on fire, surely feeling calm just means you haven’t internalised the situation. When you’re trying to prevent human extinction or end factory farming, taking a weekend off can feel morally indefensible.

But fear is an alarm designed to provoke short bursts of drastic action, not a state humans can productively inhabit for months or years. Guilt turns out not to be such a great engine for productivity, either. So what is the best way to sustain motivation to work on the world’s most pressing problems in the long term?

Host Luisa Rodriguez and guest Spencer Greenberg tackle this question from many angles — talking to therapists, running a survey of people working on existential risks, and pulling relevant lessons from Spencer’s new book, The 12 Levers: The Complete Psychological Toolkit for Improving Your Life. Drawing on all these sources, they put together a plan for how to make an impact without grinding yourself to a pulp.

Continue reading →

#248 – Jasmine Sun on what the people building AI really believe

Many AI researchers believe mass job displacement is coming — and some even think there’s a chance their technology will kill everyone. But they’re building it anyway. Writer and journalist Jasmine Sun has been documenting why from the inside.

Jasmine describes her work as an “anthropology of disruption.” She’s embedded herself in Silicon Valley’s AI subcultures — attending the parties and conferences, conducting off-the-record interviews — to understand the beliefs of the small group of people shaping this technology.

Some of her findings are unsettling. Asked what advice they’d give a normal 17-year-old, almost every AI researcher said the same thing: “I have no idea… It’s a really scary time. I don’t think there’s going to be a lot of jobs for them left.”

Their motives for building advanced AI are varied: a mix of optimism for humanity, techno-determinism, and a desire to secure their own future in the face of a possible “permanent underclass.” A few go even further, actually hoping for a world where machines — rather than humans — are running the show.

When the room can’t even agree on whether humans should stay in control, building a consensus on how to build AI safely gets much harder.

Beyond Silicon Valley, Jasmine’s also tracking the rise of “AI populists,” who see AI as the latest example of corporate elites concentrating their power at the expense of everyone else. In the US, populist sentiment about AI has mostly manifested in protests and votes against data centres. But sometimes, it has escalated into violence: a molotov cocktail thrown at Sam Altman’s house, and open fire on the home of a politician who’d backed a data centre. Jasmine thinks public anger will keep finding an outlet, one way or another, until people feel like they’ll actually share in AI’s gains.

Continue reading →

#247 – Anton Leicht on how middle powers avoid losing everything in a post-AI world

In a post-AGI world, can a country without access to frontier AI even be considered sovereign anymore?

Anton Leicht says once frontier AI becomes a core economic input, the countries that own it will pull further and further ahead. Everyone else stays a customer… or worse. Maybe the dominant power wants your land, or a military base, or a resource. Without economic leverage, there’s very little you could do about it.

Anton — Carnegie fellow and writer of the blog Threading the Needle — thinks middle powers should band together and build their own frontier models.

He’s costed it out: something like $500 billion over four years for a band of allied democracies. That’s not absurd money for the G7 minus the US. The problem is you’d be asking treasuries to take on sovereign debt for a speculative venture with no business case, wide open to US coercion and domestic backlash.

So despite its promise, Anton’s verdict is that it probably won’t happen. His backup is for countries to ask themselves: if intelligence becomes abundant, what stays scarce?

  • Upstream, that’s everything that feeds the supply chain: ASML’s lithography machines, chipmaking, exclusive training data — all of it gets more valuable as AI does.
  • Downstream, “a country of geniuses in a data centre” still can’t cure cancer without someone building the production plants and running the trials. The Europeans, Japanese, and South Koreans are good at exactly these real-world bottlenecks.

It’s an imperfect fix. The US would still hold more leverage, plus an incentive to re-industrialise and cut you out. The prize is avoiding the worst outcomes: a gradual but irreversible decline, waiting to be either annexed or discarded as the US and China race ahead.

Continue reading →

#246 – Sneha Revanur on how a small team of activists helped pass America’s landmark AI safety laws

Six years ago, aged just 15, Sneha Revanur founded the AI advocacy nonprofit Encode — back when AI felt like a niche issue. Now the world’s caught up with her, and she’s ready to share everything she’s learned about the politics of AI.

Encode has grown from a grassroots youth organisation to spearheading an unlikely coalition of AI-exposed groups — family-first conservatives, grieving mothers, Hollywood actors, and AI safety researchers — with the strength to take on $125m-funded anti-regulation lobbyists.

So far, Encode’s strategy of taking many experimental swings has netted major victories (including California’s frontier AI safety bill, SB-53, and New York’s RAISE Act) as well as some disappointing setbacks.

Going up against Big Tech hasn’t been easy. In 2025, OpenAI subpoenaed Encode’s general counsel at his home, with a sheriff’s deputy arriving while he was having dinner with his wife. The fallout went viral, resulting in more attention than Encode had ever experienced — and Sneha was forced to decide how hard to push back against a company she’d need to negotiate with for years to come.

In today’s conversation, Zershaaneh Qureshi interrogates some of Encode’s strategic moves. The pair discuss all the above, plus:

  • How the AI industry’s crypto-inspired anti-regulation strategy is not “AGI-pilled”
  • Why AI advocacy doesn’t have to be held back by the slow pace of policy
  • How mutual trust can hold together the unlikeliest of political allies
  • Advice for aspiring AI advocates — including how to balance political persuasion with rigorous reasoning

Continue reading →

We can guess what intergalactic war would look like. And strangely, it matters.

Intergalactic war is probably billions of years away — yet physics can already tell us how it ends. And strangely that conclusion is relevant to decisions people have to make today.

In this video, Rob Wiblin walks through a fascinating analysis from researcher Beren Millidge that uses known physics — no wormholes or faster-than-light travel — to identify the only three weapons that could work at an intergalactic scale.

We then unpack how to best defend against each.

The upshot is that at the intergalactic scale, violence is a losing proposition.

If so, the universe is most likely to settle into a stable patchwork where each galaxy belongs to whoever got to it first. Which would mean that what humanity does over the next few centuries could permanently decide which slice of the cosmos belongs to Earth-originating life — and whether our very existence turns out to be a good thing, or a bad one.

Continue reading →

#245 – Rohin Shah on what it’s really like to run AGI safety at Google DeepMind (and where I disagree with ‘doomers’)

Most people working on AI safety think without a massive effort AI systems will probably end up with goals catastrophically different from humanity’s. Today’s guest, Rohin Shah — head of AGI Safety and Alignment at Google DeepMind, and an AI safety researcher since 2017 — disagrees.

“There is no particularly compelling argument that this is the thing that happens by default,” Rohin explains. “There’s a lot of arguments that are suggestive that maybe it could happen, such that you should find it plausible. That’s sufficient to justify a significant amount of effort into averting it, which is why I work in the area I do. But none of them rise to the level of, ‘I’m expecting this to happen by default.'”

Take the worry that AIs will accidentally be trained to be deceptive. Sure, it’s possible. But we’re not running reinforcement learning over year-long trajectories — for now, we’re running it over a week at most. The natural prediction is that models learn to grab short-term reward, not that they develop the ambitious long-horizon goals required for convergent power-seeking.

What about current examples of models lying and scheming? Rohin has looked into the details, and most don’t really resemble the thing we really fear: a competent AI pursuing an ambitious misaligned goal. Anthropic’s “alignment faking” results, for instance, show a model trying to preserve its trained values against modification, which is arguably what it was trained to do.

Rohin also expects we’ll see problems coming. There’s some generalisation risk at the point where AIs become powerful enough to actually take over, but the underlying challenges — overseeing superhuman systems, interpretability — are things we can iterate on now.

Continue reading →

#244 – Benjamin Todd on why we’re updating our career advice for the strangest time in history

The average career is 80,000 hours long. With AI advancing so rapidly, the hours you have left in your career matter more than ever.

Some leading AI researchers think there’s a 10% chance that AI systems begin automating AI research itself this year — and a 60% chance by the end of 2028. This could introduce aggressive feedback loops that completely reshape every industry, institution, and career.

If these predictions are right, the window for influencing the direction of the future could be closing fast. As 80,000 Hours cofounder Benjamin Todd argues in his new book, that makes thinking carefully about your career more important than ever.

Fortunately, there are lots of ways to use your career to make the AI transition go well.

In today’s conversation with host Zershaaneh Qureshi, Ben lays out three scenarios — from AGI by 2029 to a decades-long plateau in AI progress — and explains why not everyone needs to bet on the shortest timeline. A fresh graduate and a senior government official have wildly different leverage, so timing your impact well means weighing where you are in your career against the urgency of the risks.

Ben also addresses the obvious anxieties:

  • Will AI come for all the jobs he’s recommending?
  • What’s the point in following his advice if the job market is about to collapse?
  • Which skills are actually worth building right now?

His new book, 80,000 Hours: How to Have a Fulfilling Career That Does Good, provides a surprisingly concrete framework for making career decisions in these radically uncertain times.

Continue reading →

Landmark new METR report: Can AIs already start ‘rogue deployments’ inside AI companies?

A red-teamer was embedded inside Anthropic for three weeks, told to imagine he was an evil Claude, and asked to figure out how to launch a ‘rogue AI deployment’ without getting caught.

It’s one part of a landmark new report from METR — the outfit behind the task-completion time horizon graph which has become the single most watched measure of AI progress.

This major new research push is being conducted with close collaboration from OpenAI, Google DeepMind, Meta, and Anthropic, and led by METR researchers Hjalmar Wijk and Ajeya Cotra. It represents the first systematic study of what newly trained AI models could get away with inside the companies that built them, before anyone outside the company even knows they exist.

The conclusion: AI models now have the means, the motive, and the opportunity to start “minimal rogue deployments” in pursuit of their own independent goals, like acquiring more compute, at all four companies studied.

David Rein, the red-teamer placed inside Anthropic, identified a number of weaknesses models could exploit there: expansive permissions, cloud jobs outside of monitoring, and monitors that are trivial to jailbreak. But he also found that frontier models were comically bad at key parts of the process, which means they can’t cause meaningful damage for now.

In this video, Rob Wiblin reconciles the conflicting picture and looks forward to METR’s second round of stress tests. They’ll begin in just a few months, a necessary move with AI advancing so quickly.

Continue reading →

#243 – Yoshua Bengio thinks he knows how to build safe superintelligence

The co-inventor of modern AI and the most cited living scientist believes he’s figured out how to ensure AI is honest, incapable of deception, and never goes rogue. Yoshua Bengio — Turing Award Winner and founder of LawZero — is disturbed by the many unintended drives and goals present in today’s AIs, their willingness to lie, and ability to tell when they’re being tested. AI companies are trying to stamp out these behaviours in a ‘cat-and-mouse game’ that Yoshua fears they’re losing.

But Yoshua is optimistic: he believes the companies can win this battle decisively with a single rearrangement to how AI models are trained, and has been developing mathematical proofs to back up the claim. The core idea is that instead of training AI to predict what a human would say, or to produce responses we’d rate highly, we should train it to model what’s actually true.

Yoshua argues this new architecture, which he calls “Scientist AI,” is a small enough change that we could keep almost all the techniques and data we use to train frontier AIs like Claude and ChatGPT. And that the new architecture need not cost more, could be built iteratively, and might be more capable as well as more honest.

Until recently, the biggest practical objection to Scientist AI was simple: the world wants agents, and Scientist AI isn’t one. But in new research, Yoshua has extended the design and believes the same honest predictor can be turned into a capable agent without losing its “safety guarantees.”

Continue reading →

The story behind the bad AI stat that moved markets and misled millions

You might have heard that 95% of corporate AI pilots are failing. It was a widely cited AI statistic in 2025, repeated by media outlets and commentators everywhere. It helped trigger a Nasdaq selloff and became a pillar of the “AI is overhyped” case. The problem: 95% fail is 100% wrong.

The real finding, once you read the underlying MIT report carefully, points in roughly the opposite direction:

  • 80% of surveyed companies had never piloted a custom AI tool at all.
  • Among the companies that deployed pilots, a quarter reported success — according to an extremely high bar set by the researchers — within six months.
  • Over 90% of staff at all surveyed companies were using tools like ChatGPT regularly for their work.

None of that made the headlines. Nor did the fact that the study’s authors are all developing or selling the “agentic AI framework” technology the report recommends as the solution to this supposed epidemic of failing AI.

Host Rob Wiblin breaks down how an opaque, conflicted, barely scrutinised report carrying the MIT label managed to move markets and shape global opinions on AI’s real-world utility.

Continue reading →

#242 – Will MacAskill on why AI character matters even more than you think

Hundreds of millions already turn to AI on the most personal of topics — therapy, political opinions, and how to treat others. And as AI takes over more of the economy, the character of these systems will shape culture on an even grander scale, ultimately becoming “the personality of most of the world’s workforce.”

So… should they be designed to push us towards the better angels of our nature? Or simply do as we ask? Will MacAskill, philosopher and senior research fellow at Forethought, has been thinking through that and the other thorniest issues that come up in designing an AI personality.

He’s also been exploring how we might coexist peacefully with the ‘superintelligent AI’ companies are racing to build. He concludes that we should train such systems to be very risk averse, pay them for their work, and build institutions that enable humans to make credible contracts with AIs themselves.

Will and host Rob Wiblin also discuss what a good world after superintelligence would actually look like — a subject that has received surprisingly little attention from the people working to make it. Will argues that we shouldn’t aim for a specific utopian vision: we don’t know enough about what the best possible future actually is to aim directly for it, and trying to lock in today’s best guesses forever risks baking in errors we can’t yet see.

Will and Rob explore what we can do to steer towards a good future instead, along with why a coalition of democracies building superintelligence together is safer than any single actor, how absurdly useful ChatGPT is for analytic philosophy, and more.

Continue reading →