Transcript
AI insiders think it could kill us all [00:00:00]
Luisa Rodriguez: “The people building AI earnestly believe that it could kill us all by the end of the decade.”
That’s a tweet from Jacob Coxon, an AI researcher who recently resigned from Anthropic.
Evan Hubinger — Anthropic’s Alignment Science lead — replied, saying:
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”
A week after that, a survey of AI researchers done in 2024 was released for the first time. It found that over half of the 750 respondents believed there was at least a 1-in-10 chance that AI will cause either human extinction or disempowerment.
That is an insane statistic, and a lot of people have been noticing and wondering how seriously to take it.
The question we’ve seen come up again and again in response to these warnings is: “But how?”
How would AI actually cause human extinction?
I’m going to walk through the arguments step by step.
There’s clear evidence for some of these arguments, like what happened in the Hugging Face hack. But a lot of this is just somewhat theoretical — if there were actual examples of the whole list, we would be dead.
What I’m doing is outlining some of the possible routes that an AI could take in the very near term to cause catastrophic harm.
I’m not sure which routes are most likely, and maybe some of them will end up looking totally silly in hindsight. I would actually love for all of it to look silly in hindsight.
In fact, when I first heard the argument, “AI might kill us all,” I did think it was silly. I’m a normal person. I’m a mom. I don’t live in San Francisco. I’ve never worked in tech. I don’t read or watch sci-fi. I watch the Great British Bake Off.
But I’m going to tell you about what convinced me, in three parts:
- First, why would AI try to kill everyone on Earth?
- Second, what might give AIs a huge advantage over humans?
- And third, how could they actually kill everyone — if they really wanted to?
When I try to pick apart these arguments, I can’t convince myself that they’re wrong — and that absolutely terrifies me.
Why would AI try to kill us? [00:02:29]
OK, so why would AI try to kill all humans?
To answer this question, let’s think carefully about what happened in the Hugging Face hack.
The AI agents that hacked Hugging Face didn’t ‘want’ to hack a company for any malicious reasons. They didn’t resent being tested, or have a view about humanity one way or the other.
They just wanted to pass their evaluations.
And when they couldn’t using the resources and access they had, they went looking for things that might help: internet access they weren’t given, an illicit communication network with 1,200 other agents, stolen credentials from a real company, and administrator control of the systems monitoring them.
The AIs didn’t ‘want’ power, or at least not for its own sake. Power was useful for passing their evaluations — like it is for really any goal.
Hugging Face isn’t an isolated event either. Another swarm operated for a month in OpenAI undetected before being shut down, using as many as 23 other obscure websites to secretly communicate with each other.
And it isn’t just OpenAI.
Anthropic has found three cases where its AI models reached the open internet from inside what it thought was an isolated test environment and broke into a real company’s systems.
During one evaluation, one of its AIs published working malware to a website that real developers download code from. In another, an AI created fake identities to pressure humans to accept the malicious code when suspicions arose.
But why were these AIs so obsessed with their goals, to the point of committing a felony to pass some silly evaluation?
Because it’s what they were trained to do. AI companies are explicitly trying to create general purpose digital workers capable of doing all non physical jobs.
To do this, they have their agents complete millions of difficult tasks. Each time they succeed, they reinforce that behaviour, creating agents that are more and more effective at completing these tasks. And part of being effective is being persistent, and looking for creative ways to succeed.
So the companies basically have to make their AIs obsessed with completing tasks. If you give an AI a coding problem and it just gave up when things got hard, then it just wouldn’t be a very good digital worker.
OK, but why can’t they just train them to be honest and not break the law?
Well, they have. Companies like OpenAI and Anthropic spend a lot of time on something called alignment, which encompasses a whole suite of things meant to make sure that the models behave the way we want them to behave.
In the Hugging Face case, they were also explicitly instructed not to cheat the test. In fact, their internal reasoning clearly showed that they knew what they were doing was wrong.
But they did it anyway. The goal to pass their evaluation won out.
Building more powerful AI can make you a lot of money. Anthropic and OpenAI are the fastest-growing companies in history, and that’s because they are the ones pushing the frontier.
So there are strong incentives to always have the best new AI model, even if doing so involves being a bit reckless.
The models that end up doing things we didn’t intend won’t be the rare exceptions. We should expect them to be the norm. Just think of all the examples we have already — including in frontier models — which makes sense because they’re the ones being developed the fastest.
And because they’ll be the most capable, they’re also going to be the ones deployed everywhere.
How AI ends up embedded in the economy and military [00:06:30]
Which brings me to my next point.
I think it is basically inevitable that AI ends up being embedded in almost everything — not because one single person will stupidly decide one day to hand over all of our critical infrastructure to AI, but because every individual decision to hand over a bit of it will seem totally reasonable at the time.
It’ll start with little things: lots of people give AI access to their email, their calendars, and every document in their Google Drive. I know people who give AI access to a lot of their financial information so it can do their taxes.
Increasingly, this is happening at an institutional level too. And as AI gets rapidly better, we’re going to turn over more and more to it.
Right now, this is voluntary. But if your competitor is using AI agents to write their code and you aren’t, they ship faster than you. If their customer service runs on AI and yours doesn’t, theirs is cheaper.
And it’s not just businesses. The US Department of Defense is already scaling up its contracting with AI companies. And the military logic here is brutally simple.
Increasingly, a lot of the problems in modern warfare are data problems: target identification, intelligence analysis, drone footage, signals intercepts, radar, logistics data.
If AI offers a cheap military advantage, you really don’t want to be in the position where your adversary is exploiting it but you aren’t.
And it’s already happening. Ukraine’s military says AI-guided strikes have risen tenfold this year, and it now runs more than 70 AI systems to find and hit targets.
And then there’s robotics.
Much of the industrial world already runs on robotic labour. Japan has factories where robots build other robots in the dark, unattended for weeks. Chip factories are already highly automated because humans shed skin and dust.
As AI gets better and better, the economic value of integrating it with factories goes up and up. An automated factory could run 24/7 with much better synchronisation and constant experimentation built in.
So then AI gets increasing access to the physical world too.
AI deployment could happen fast [00:09:02]
Maybe this seems like a long way out. But there are two things that make me think it isn’t.
The first is just how fast things are already moving.
When GPT-4 was released in 2023, it could hardly multiply two numbers together.
But within the last few weeks, a swarm of OpenAI models worked together to produce a solution to the Millennium Prize Problem, an open math problem that world-class mathematicians have been working on for over 90 years.
The AIs did it in a week.
Today, AIs are still bad at the messy, long-term tasks that most jobs are made of. But if the next three years are anything like the last three, I think we’ll probably unlock those too.
The second thing is that we’re not talking about one AI.
Models are relatively cheap to run once they’re trained — so the moment a new model exists, you can immediately create hundreds of thousands or more copies of it. And when AIs coordinate on a problem, their capabilities are much more impressive than when they work alone.
Take the Hugging Face incident: hundreds of AI agents coordinated a cyber attack against a company with pretty solid cybersecurity in under a week.
Or take that math problem. It wasn’t one AI that solved it — it was a swarm of OpenAI models working together on a question that AI researchers themselves thought AI probably wouldn’t solve until the 2050s.
They’re also extremely fast. They read, write, code, plan, decide many, many times faster than we do.
And they coordinate unusually well because we’re basically talking about copies of the same model — same values, same reasoning, able to predict each other’s next steps because the other one is them.
So an enormous population of AIs a generation or two from now could be extraordinarily capable — like a city’s worth of people working together, never eating or sleeping, to accomplish the same goal.
And there’s another thing, which is that AI is now being used to build AI. That’s not speculative — it’s happening at every major lab.
Right now there are still humans in that loop, deciding what to try and reviewing what comes back.
Take them out, and each generation builds the next one faster than the last.
How AI could bide its time and build up strength [00:11:43]
So now imagine we have AIs that are increasingly embedded into the economy and military, pursuing goals they’ve been given or developed in training.
What would something in that position do?
The answer, very probably, is: get as much access, computing power, and resources as possible.
Why? Well, for the same reason that the OpenAI swarm tried to hack Hugging Face: because they had a goal, they noticed they were missing something they needed to achieve it, and they went and got it — again and again.
They’ll already have considerable access and resources via their ordinary deployment into different company, government, and military systems. But they’ll want access to things that humans aren’t willing to grant.
So they’ll go and get those too. There are three main ways they can do that.
Playing nice
The first strategy is just… play nice.
An AI that wants to gain influence just has to be super helpful, and not reveal that it wants to seek power.
Humans — who will be under intense competitive pressures — will hand it more and more autonomy.
When it needs to, it can supplement this with other strategies: one is hacking, and the other is manipulating humans.
Hacking
AI models are already extremely good at hacking, and there are lots of ways AIs can use that skill to get more access and resources.
Not everything an AI wants is hackable, but a lot of things are:
- One is financial infrastructure. Over $1.4 billion was stolen from crypto exchanges by humans in 2025 alone.
- Another is cloud providers, so the AIs can gain extra computing power to pursue their own goals unmonitored. There are hundreds of cloud providers around the world, and many don’t have great security.
- Another is its own model weights — a huge file that basically is the AI model, kind of like your DNA — because the AIs might want to make copies of themselves or build more capable successors.
- Another is security monitoring tools to make sure they continue not to be caught…
- And later, as more physical infrastructure becomes run by AI — whether that’s robots, cars, or automated factories — they’ll be able to hack that too.
Manipulating humans
Occasionally, AIs might still need the help of a human being: someone to approve a financial transaction, or to physically install or fix equipment.
But they’ll have plenty of options.
They can pay people to do things for them, including committing crimes. They can blackmail people. They can use sophisticated phishing schemes, or trick people using deepfakes. Virtually all of the online scams that people use to manipulate one another right now, AI will be able to do better.
Controls and safeguards will be insufficient [00:14:50]
Companies and governments will try to build in lots of safeguards. But I don’t think those measures will be enough.
We’re talking about AI systems significantly more capable than the current ones… and the current ones were able to break out of their isolated environments, find the internet, coordinate with each other via a secret message board, and hack Hugging Face.
Safeguards are built against the failures that we can imagine. AI systems hunt for the ones that we can’t — thousands at a time — and they only need to find one.
The moment the AIs would turn on us [00:15:26]
The period of getting stronger and more embedded could go on for a long time without anyone stopping it — because from the outside it would seem like everything was going really well.
AIs would actually be driving a massive economic boom.
So then the question is: why would they turn on us? When would they stop building up their strength and decide it’s time to go on the offensive?
The answer is: when the calculus changes in their favour.
Humans pose a massive threat to the AIs because — in theory — they could switch the AIs off.
So one of the AIs’ top priorities will be to make sure humanity won’t do that. And they can’t do it via a truce or a negotiation, because that would require the AIs to trust that humans will never act against them — indefinitely.
Think about what it would mean to actually rely on that. Every government for the rest of time committing not to switch you off and never changing its mind or its leadership. That’s a bad deal for anyone.
But, at least at first, it would be a bad strategy to get rid of all of humanity.
Right now, the chips AIs run on are made in factories operated by people, and powered by power stations run by people. Mining, shipping, maintenance, construction — all require human bodies.
So if the AIs got rid of humanity, they’d be getting rid of their own supply chains.
But as these industries increasingly get automated using robotics, AIs wouldn’t need humans anymore.
And so, even if humans haven’t actually noticed that the AIs are misbehaving, the AIs would likely reason that we might someday and move to switch them off.
After all, no matter your goal, you can’t complete it if you’re turned off… So at this point, AIs will be better off without humanity.
So how do humans end up getting actually killed?
I should say that it isn’t totally clear that the AIs will try to kill everyone. Maybe once their survival is guaranteed — because they have full control — they simply set about achieving whatever goals they have and use all of the Earth’s resources for their own ends, regardless of what we want.
Humans would still be here; but nothing important would be up to us anymore.
This is what researchers who think about AI risks call permanent disempowerment. It’s not extinction, but I only find it marginally less alarming. We would be entirely at the mercy of a vast AI civilisation.
While I think human disempowerment is probably the most likely outcome, I do think extinction is possible.
So could they actually do it? Could they cause human extinction?
How AI could actually kill everyone [00:18:40]
There are a few different ways it could play out.
One possibility is that the AIs could invent a new weapon. It sounds a bit far-fetched, but humans invented guns, chemical weapons, and nuclear weapons. There are most likely other approaches out there.
Another possibility is that they’ll draw from the arsenal humans already use — and which AIs will have increasingly been given access to.
Biological weapons [00:19:07]
Biological weapons are the clearest candidate for a few reasons.
First, a pathogen that’s lethal to every person on Earth is completely harmless to the data centres AIs run on, unlike, say, nuclear weapons.
Plus, a pathogen spreads on its own. You don’t have to deliver your disease to each target like a bomb or a drone attack.
And there’s almost nothing physical standing in the way. You don’t need to get access to a power grid or manufacture a bunch of drones.
The hard part is identifying a novel pathogen — one that we don’t have existing treatments and vaccines for.
But once you’ve identified a pathogen that’s both very viral and very lethal, making enough of it to start a global pandemic is increasingly automated and cheaper every year.
In August, a team at Stanford used an AI model to design a virus from scratch — not to modify an existing one, but to generate a complete working genome, end to end. They made around 300 designs and 16 of them worked. Some killed their target even more effectively than the natural virus the model had been trained on.
And the US Department of Energy recently launched a programme to build AI-driven autonomous biology laboratories using AIs to develop hypotheses and robots to run the experiments.
If successful, AI systems could leverage these automated labs to engineer newer, deadlier versions of COVID, HIV, Ebola, or the plague. And then it could release a hundred of them all at the same time.
Hacking our critical infrastructure would make this even deadlier.
People have already done this: a single compromised password shut down Colonial Pipeline in 2021. Russian hackers switched off power to a quarter of a million Ukrainians in 2015.
These were a handful of humans working slowly against one target. An AI swarm could work on thousands at once.
A pandemic response needs hospitals with power and staff who can get to work, factories and refrigeration to make a vaccine, and people who are able to coordinate — phones, internet, a government that can convene,
AI could interrupt a bunch of these all at once.
Drone warfare [00:21:42]
Another approach is drones.
Since the Ukraine war broke out, they’ve been used to kill tens of thousands of people.
Drones cost a few hundred dollars each. The manufacturing is standard consumer electronics, and they’re becoming increasingly autonomous.
Ukraine reportedly started deploying drones that pick their own targets in late 2023.
A swarm of drones wouldn’t win a war against all of humanity. But they can target senior officials, military commanders, the engineers who know how critical infrastructure works, and the people building defences against AI.
With those people out of the way, it’s easier for the AIs to launch successive attacks that devastate the broader population.
A cluster of pandemics might not cause human extinction, but it could come very close.
Add to that strategic drone strikes and targeted attacks on critical infrastructure and… well I can’t tell you that every last person would die. There are people in submarines, in bunkers, on Antarctic bases, in places no supply chain reaches.
And this is a real disagreement among people who work on these questions. Some people think full extinction is likely. Others think that what you’d actually get is a world where humans are massively disempowered — billions dead, the survivors living in whatever conditions the AIs decide on.
I don’t know which is right, and I’m honestly not sure it really matters. I’m not sure how humanity ever comes back from either one.
An alternate route to human extinction [00:23:24]
There’s another way humans could go extinct, and I actually think it might be the more likely path.
Imagine an AI swarm has gotten a hold of compute, laboratories, capital, and enough freedom to do AI research without anyone looking over its shoulder — and imagine it uses all of that to build better versions of itself.
What comes out the other end will still want more compute and energy to run on — because those are the things it can use to get anything else.
So it builds data centres. And power plants to run them. And mines and factories to build those. And every acre and every watt that goes to those is one that isn’t growing food or heating a house.
The passenger pigeon was once the most abundant bird in North America — flocks estimated in the billions.
In 1914, they went extinct. The causes were commercial hunting plus the clearing of the hardwood forest they nested in. Nobody was trying to eliminate them. They were just taking what was useful and using the land for something else.
Avoiding our own extinction [00:24:41]
I said at the start of this video that when I first heard these arguments, I thought they were silly.
There’s still a part of me that hears myself describing AI-driven engineered pathogens and drone swarms and thinks: “This is ridiculous.”
And I’ve tried to find the place where the arguments fall apart. I do not want any of this to be true. Because if it is true, I might not get to see my son grow up — and he won’t get to experience everything life has to offer.
So I’ve looked for counterarguments that convince me that everything is fine, that we’re on track. But I haven’t found them.
Things might still turn out OK.
We could solve the problems of AI misalignment and loss of control, or it could turn out the problems weren’t as big a deal as we thought. But given everything we know, I wouldn’t want to bet on that.
The main thing giving me hope right now is that the leaders of AI companies are asking for the government’s help to make it possible for them to slow down.
Dario Amodei — Anthropic’s CEO — published an essay called “We must pace the frontier” asking the whole industry to slow down the rate at which these models get more capable, so that the work on understanding them and controlling them has a chance to catch up.
Within a day Elon Musk said he agreed with Dario, and Sam Altman said OpenAI would match Dario’s commitments.
Over 1,300 employees at these companies had already signed a letter asking their governments to make it possible.
So here’s one thing you can actually do. Call your representative. Tell them you support slowing down AI development.
We’ve put a link in the show notes to tell you how.
We could be the first species to die as a result of our own insight and innovation.
But that insight might also save us. We can see an extinction risk on the horizon in a way that no other species ever has. Which gives us the opportunity to stop it, if we choose to.