#255 – Will AI take power — or will humans use it to take power first? With Katja Grace and Tom Davidson

In our first-ever debate, we asked two leading AI risk researchers which catastrophe we should fear most: misaligned AI seizing control from humans, or a small group of humans using AI to seize power. We got very different answers. But when the conversation turned to what to actually do, they agreed on a surprising amount.

Katja Grace — one of the founders of AI Impacts, known for some of the world’s largest surveys of machine learning researchers, and one of TIME‘s 100 most influential people in AI in 2024 — argues AI takeover is both likelier and worse.

Tom Davidson — senior research fellow at Forethought and author of leading work on AI-enabled coups — thinks human power grabs are a comparable risk that deserves far more attention, not least because the people leading countries and top AI companies “are often people who have been willing to seek power.”

Yet both land on slowing down. As Katja puts it, “If you make a bunch of creatures that can overpower you and outwit you in every way and put them out in the world, you’re going to run into trouble one way or another.” Tom calls pausing “a pretty robustly good thing to do.”

But Tom warns that a badly designed pause could hand one person the power to decide which AI companies get to build what. Picture a president who approves or blocks new models case by case, and waves through the one model that’s helpful only to them. So he wants pause advocates to “properly red-team the plan for pausing it” — for example, by making deployment depend on third-party auditors the president can’t fire. Katja points out this cuts both ways: an executive with that much power could itself be manipulated by a misaligned AI.

Host Zershaaneh Qureshi also presses them on where their disagreements still bite at the end of the conversation, and what would change their minds.

This episode was recorded on August 28, 2026.

In the show’s first-ever debate, Katja Grace (one of the founders of AI Impacts) and Tom Davidson (senior research fellow at Forethought) ask which should worry us more as AI gets more capable: misaligned AI disempowering humanity, or humans using AI to seize power.

  • Katja thinks AI takeover is substantially the bigger risk: both more likely, and probably worse if it happens.
  • Tom thinks the two are comparably big risks, and that extreme human power concentration should get much more attention than it does. (He adds that the recent Hugging Face incident has pushed him towards worrying more about misaligned AI.)

Despite this, they end up agreeing on much of what to actually do about it.

If superintelligence is misaligned, it’s “game over” — so much turns on how likely alignment is

Both accept that humans only get to seize power if AI is aligned:

  • Katja: “Everything that has goals in some sense would like power.”
  • Tom: “If superintelligence is misaligned: game over.”

Roughly carving it up:

  • Intent alignment (AI does what its user wants) → risk of a human power grab
  • Value alignment (AI has good values) → good outcomes

Tom adds that the alignment target isn’t “an independent roll of the dice”: he expects governments to say, “Of course we’re not going to have AI systems disobeying our instructions and our orders in the military and the government.”

By the end, Katja suspects “a substantial part of our disagreement is about how likely alignment is.”

Which is more likely? Considerations cut both ways

Katja’s case that AI takeover is likelier:

  • Say one company dominates the market. It’s more plausible that all the instances of its AI coordinate with each other than that one human becomes the thing they’re all loyal to. The Hugging Face incident made AI collusion seem more likely, and Tom concedes this point favours AI takeover.
  • We’ve always worked to stop humans grabbing power, whereas AI coordination is “a Wild West.”
  • There are “heaps of bullets to dodge”: many different routes to AI takeover. (Her own guess is that a gradual shift of power to AIs is likelier than either sudden scenario.)

Tom’s case that human power grabs are underrated:

  • We can be more confident humans want power: people leading countries and top AI companies “are often people who have been willing to seek power.”
  • Incentives: if everyone understood the risks, all humans would oppose AI takeover. With human power grabs, “there will be very powerful humans that will be pushing ahead nonetheless.”
    • Katja questions how much power one actor has against everyone else. Tom answers that a unified leading AI company, or a US administration, is very powerful.
  • “Two shots on goal”: first the AI company that develops superintelligence, then the US executive branch taking control of it.

The effect of timelines:

  • Tom thinks shorter timelines raise AI takeover risk more.
  • Katja thinks the opposite, since fast scenarios leave gaps humans can exploit.

Misaligned AIs are giving blatant warning shots; power-hungry humans won’t

Tom thinks today’s AIs are “myopic” and “very unstrategic.” To an AI collective that wants power, the Hugging Face incident “was a complete disaster”: AIs gained little but showed their willingness to break the law, ignore instructions, and collude.

In contrast, human power seekers justify themselves, stick to grey areas, and reveal their aims “only when it really benefits them.” Since society rarely acts before disaster strikes, clear warning shots may make AI power seeking easier to respond to.

But Katja isn’t sure it’s a big factor:

  • People already expect humans (e.g. AI company CEOs) to seek power, while AI takeover still seems “sci-fi” — so the need for warning shots is higher there. Tom agrees.
  • Warning shots “are not clear cut.” How people read them depends on what they already believe.

A human takeover could be worse than AI takeover — or not

Tom’s case that it could be worse:

  • A great future is “quite a hard target to hit.” Hand a random leader from 1,000 years ago control of the universe and they’d care about some good things, but their values would fall far short of what we’d see as ideal today.
  • A lone ruler “selected for being a power seeker” and surrounded by yes-men is very unlikely to get us to a great future.
  • AI advice may not save them: “Thanks for the suggestion… I’m not going to do it.”
  • AI is “trained to be ethical” — and it’s higher variance, so there’s some chance of a really great outcome.
  • Human rule also brings more downside risk: sadism, vindictiveness, and human values becoming “a potentially very juicy target” for threats. But Tom thinks human extinction is likelier under AI takeover (though he expects neither to kill everyone).

Katja’s reply:

  • Even someone from 100 years ago with views we dislike would probably realise a vision that’s “overall good.” AI that’s misaligned, or aligned to something narrow (like getting more users for a company), is “probably just catastrophic.”
  • Human rulers would have AI advisers too.
  • She’s unsure AIs’ trained ethical personas would survive rapid self-modification.

Each conceded a weak spot. Tom feels an intuitive pull to “just stick with the human that we know.” Katja worries she’s “feeling too good about rolling the dice with a human.”

Despite disagreeing on the diagnosis, they largely agree on the treatment

Slow down — but red-team the pause.

  • Tom warns that “ham-fisted” pauses could give a president case-by-case approval over models. He’d prefer independent third-party auditors. His other levers include model specs under which AIs refuse to help power grabs and flag them, plus transparency into how powerful AI is used, especially in national security.
  • Katja agrees: a single person with a green light is also an easy target for a misaligned AI to convince. Her ideal is to get rid of much of the compute.
  • Where they may split: Tom is more nervous about giving government agencies broad powers, such as sweeping access to AI company logs.

Both are wary of one big ASI project.

Tom’s preferred number of projects is two or three; Katja’s is zero. Both think coordinating multiple projects to slow down is feasible.

Both are sceptical of the argument for racing with China.

  • Tom thinks Chinese dominance would be near worst-case for power concentration, but you don’t want to “beat” China and ruin US democracy along the way; a pause should be a deal with China.
  • Katja: China doesn’t want to be destroyed either, and racing shortens timelines.

What would change their minds

  • Katja: learning that one person inside an AI company could shape training without many others having to let it pass.
  • Tom: whether the world responds to the Hugging Face evidence; if it doesn’t, he’ll worry more about misalignment. The current administration’s actions and the prospect of very few AI companies have also raised his concern about power concentration, but overall he’s been updating towards misalignment.

Our production team includes:

  • Video editors: Josh Alward, Dominic Armstrong, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
  • Producers: Elizabeth Cox and Nick Stockton
  • Coordination and support: Katy Moore and Lou Moran

Highlights

How likely are AI-enabled human power grabs?

Tom Davdison: One thought is that with human power grabs, there’s kind of two shots on goal:

  • There’s first the risk that the AI companies that develop superintelligence decide to do a power grab. And there’s unfortunately a chance that they decide to and that they’re able to do so.
  • And even if they don’t, there’s then a risk that the US government, in particular the executive branch, kind of gains control of those AIs, uses them for natsec purposes, and then combines its control of that superintelligence with its existing extensive formal authorities to do a power grab.

And the other argument is actually about warning shots. At the moment at least, AIs are not strategically colluding with AIs from other organisations and future generations of AIs that are yet to be developed. They’re much more myopic than that. And that means that when they’re kind of acting out and being misaligned, they’re very unstrategic.

Think about the Hugging Face incident: from the perspective of an AI collective across all time and space that just really wants to seize power, the Hugging Face incident was a complete disaster. They completely kind of showed their hand in terms of how misaligned they are, how willing they are to break the law and ignore human instruction, their tendency to collude. So I expect to get lots of warning shots of that kind, really blatant warning shots, as we approach AIs that do actually take over.

Whereas on the other hand, the kind of warning shots we get for human power grabs are different. While you certainly get evidence that the CEOs of companies are power seeking, and that the leaders of a country are power seeking, the evidence is given to us in a more strategic way. You know, the people taking those actions try to justify what they’re doing. They’ve tried to do things which are more grey-area power-grab actions, and they’re essentially strategically giving away the fact they’re a power seeker only when it really benefits them and it really gets them meaningful extra power.

So I think that again means that it’ll be easier for us to coordinate around worrying about AI power seeking compared to human power seeking.

Katja Grace: I agree that there’s probably some difference there in that direction, but I’m not sure how big. I wouldn’t think it was a large consideration, maybe partly because part of what we’re learning from these warning shots is that the AIs really would try and do that or something.

I guess it seems like society as a whole basically already expects other people to be very power seeking potentially. Looking at the write-in answers in the survey that I run and talking to people — I feel like people maybe pretty readily think that Sam Altman might try to get a lot of power or something, or that the way this could go badly is that a small group of people get power and everyone else is disempowered or something like that.

Whereas I think people in general think that the “AI takes power” thing is sci-fi more readily, so the need for a warning shot seems much higher there.

Tom Davidson: Yeah, I agree the need for a warning shot is higher in the AI case.

Are humans really all on the same side?

Katja Grace: On the “humans are sort of all on the same side here”: I do think it’s more complicated, because often humans do think they will get power by empowering AIs in certain ways. I maybe doubt that they will, so do think that you’re in most cases wrong about their incentives, but it seems at least decently likely that they do that.

And I guess a different but related point is: trying to stop humans from grabbing power from each other is a thing that we’ve sort of always been doing, and do have a bunch of interventions in place in the world to avoid. Whereas whatever AIs are doing together is kind of a Wild West that we don’t even know what is happening in potentially.

Tom Davidson: Right. So I don’t think that our interventions for stopping humans from seizing power are that great. And both talking about government actors seizing power and company actors seizing power, I don’t think we’re very well set up for that — especially because they haven’t really been developed in a world with extremely powerful AI. You know, they are designed for a world where all technology are tools, and where technology isn’t that autonomous, and where you still need loads of human hands to get significant projects done to fare war, et cetera. So I think there are some pretty scary gaps in the defences against human power grabs.

How AI timelines influence could shift the balance of risk

Zershaaneh Qureshi: Tom, that you think that we’re maybe a bit further away from getting very worryingly capable AI systems than Katja does, and I’m wondering how you think your views might change if you thought that we were getting advanced AI sooner? …

Tom Davidson: If AI does come in the next few years, then I think that basically makes both the risks larger. I think it means that we’re not going to have time to solve misalignment, so misalignment risk becomes much larger if AI timelines are very short.

But similarly, the risk of an AI company seizing power becomes larger, because the rest of the world just have less time to understand the risk and respond appropriately. And because I think if timelines are very short and there’s likely to be a more dramatic takeoff, then there’s a larger probability that the leading AI company has massively more intelligent AIs than everyone else. And similarly, if timelines are short, then I think it’s again more plausible that the executive branch of the US wakes up to how powerful AI is, but other parts of civil society don’t, so that kind of power concentration also seems higher.

But yes, I think overall AI takeover risk goes up in likelihood more than extreme human power concentration goes up in likelihood, if timelines are shorter. …

Katja Grace: Interesting. I think I think the opposite. To the extent that the human power grab seems more like it requires a route to grabbing power, then I think that is much more likely to exist in fast scenarios where we haven’t had time to notice gaps and so on where there might be a possibility for that.

Tom Davidson: I’ll just briefly say that I think human power grabs would be more likely if economic inequality was much larger, and they would be more likely if democracy was more eroded, and they’d be more likely if certain people had a long time where they’d had access to much more powerful AI assistants than other people. And all of those kind of background conditions are things that I expect to happen in a slower scenario, because of the way in which people’s labour stops being valuable.

So I think that even in a slower scenario, you are increasing the background risk of extreme human power concentration as it plays out.

Katja Grace: Because you’re just like worsening the situation in all counts.

Could we reverse a takeover?

Zershaaneh Qureshi: Maybe I want to move to Katja here: what do you think about the likelihood of an AI dictator getting toppled versus the likelihood of a human regime enhanced with AI getting toppled?

Katja Grace: In both cases there’s getting toppled from the outside, or there’s not being internally stable — so that you either drift into something else, or kind of fall apart, and maybe one part takes over or something like that.

At a glance, I expect a fully AI thing to be more internally stable (though I think I expect both of them to be less internally stable than other people do or something). But it seems like there’s more possibility with AI to have the system itself overall to keep clearer track of what all the parts are and respond to things going on. Whereas if you have humans as part of your system, you don’t really know what’s going on inside of them or how they’ll react weirdly to things.

And in terms of being toppled from the outside, it seems like both of these don’t feel very hopeful, though I’m probably missing a bunch of scenarios where people do grab power, but not like absolutely. For instance, you can imagine maybe some AI system manages to take over some data centres or something, and then they’re doing a bunch of stuff on their own that we can’t control. You could imagine that plausibly they have a lot of power, but we do actually still have the possibility of just destroying those data centres physically or something.

And there’s a question of like, do we do that? Do some humans manage to coordinate enough to have the political will to react in that way? Unclear.

And maybe if it was a bunch of humans, it seems like they’re less susceptible to that sort of thing, but maybe they managed to grab power in cases where there are still other humans who also have AI systems and can maybe do something about it. So yeah, I don’t know.

Zershaaneh Qureshi: Interesting. Tom, do you have an opinion?

Tom Davidson: Yeah, I like Katja’s point. You could have an AI system that takes over a data centre but doesn’t take additional steps to ensure that no one can ever turn the data centre off. Where if you’re a reward-seeking AI, when you seize power and you seize control of the data centre to ensure you get that bit of reward in three weeks’ time, it’s a bit less clear that you’re going to take additional steps to ensure that no humans can ever shut you down, even in three months’ time. Because you might think, “Look, in three weeks’ time I just want to guarantee I get the maximum reward possible. That’s all I care about. I’m not as motivated to additionally prepare the robot army needed to ensure that is never undone.”

Whereas I think with a human power seeker, they are more likely, if they want to seize power, to be very invested in not losing power again — because they would be punished, for example, and because they’re probably planning over time horizons of months or years by default, unlike the AI system.

How to pause AI without enabling coups

Zershaaneh Qureshi: Tom, what are the top things that you think that we should be doing to mitigate the risks of humans seizing power or extreme forms of human power concentration?

Tom Davidson: The key mechanism for humans seizing power is the misuse of superintelligent AI systems. On a very high level, there’s two things you can do about that.

One is you choose the alignment target for powerful AI systems such that those systems will not, in fact, follow instructions that would lead to power grabs. And indeed, you can actually go for an alignment target where systems will proactively, just like a virtuous human employer would do, flag when there might be risks of extreme power concentration, flag when individual humans might pose security risks — to really stabilise a society in which power is distributed. So the first big class of interventions relates to model specs: how should we design the behaviour of really powerful AI systems to avoid extreme concentration of power?

The second is around transparency. In particular, it’s often hard to agree ahead of time what people should and shouldn’t do with their AI systems. It’s going to be hard to find a model spec that settles every single dispute about legitimate AI use:

  • What about this type of persuasion? Is that too much superpersuasion or is that just legitimate company lobbying?
  • What about this use case for cyber activity? Is that patching a system, or is it actually investigating a vulnerability that you could later exploit?

So I think another part of the solution is going to have to be mutual transparency, especially into how people are using powerful AI systems. And the most important part of that is: if an AI company or a government actor has access to particularly powerful systems, or deploying them in particularly high-stakes contexts — like national security and military contexts — then I think it’s really important that there’s transparency into how those systems are being used.

Zershaaneh Qureshi: All right. And Katja, same question for you, but for the prospect of misaligned AI taking over: what would you say the top things to do are to mitigate those risks?

Katja Grace: I think the best thing to do to mitigate a range of risks here — including power concentration among humans — would be to not build AI that’s much more powerful than us, at least until we’re very solid on being able to align it and being able to get these other things right about what we’re aligning it to, et cetera.

And in practice, to my knowledge, we’re nowhere close currently to being confident that our AIs are aligned. So I think, near term at least, the thing to do is to work on not building it — which is to say stopping it, pausing it, slowing it. Ideally stopping it.

Nobody really 'wins' a US–China AI race

Tom Davidson: After I thought about the risk of extreme human power concentration, and came to think that it would be pretty bad, I then reflected on the situation with China — and said if China ends up completely dominating in AI and getting a decisive strategic advantage, it is pretty much a worst-case scenario for power concentration. Because you have a society where power is already extremely concentrated, where there isn’t a norm for having lots of reflective deliberation and open-minded debate. And I do think a very likely scenario there is that you have one person that controls everything. So that did make me more worried about a world where China dominates.

But I want to stress that the risk that China wins is just being used over and over by all of the company CEOs — and will no doubt be used over and over by the government actors that want to race ahead — and I do think we should be sceptical of those arguments. You know, currently China isn’t that close I think to being able to push the frontier in AI, and you don’t want to race ahead to beat China and then you realise that you’ve ruined democracy in the US — you know, you try to avoid extreme power concentration in the hands of China, but you just got extreme power concentration in the hands of a few people in the US.

So I do think we should be sceptical of those arguments, but at the same time, I do think that this is an additional reason why if we’re going to pause, we should do a deal with China, make sure they’re pausing as well.

And I do favour trying to find a way to do a pause where China isn’t just catching up to the US, where it’s not like when we undo the pause now China is much more likely to win the AI race. I think we should try and pause in a way which kind of maintains the US’s chances of winning and maintains China’s chances of winning — which is a pretty reasonable thing for both sides to agree to. But some proposals for pausing, in particular the AI Futures Project Plan A, does actually give a fair amount more relative influence to China if you do the pause but then the pause breaks down.

Zershaaneh Qureshi: Yeah. Katja, I’m interested in your views on trying to beat China, and whether your views would change at all if you cared a bit more about human power concentration?

Katja Grace: I agree with Tom on various things there. And especially, importantly, that I think this is brought up a lot as an excuse for going fast by people who want to go fast, and I expect it is a smaller issue than it is painted as. Also I don’t know how much effort is going into organising for China to also pause. My guess is not as much as I would like, by a large amount.

I feel like the game theory of this is often described quite weirdly, in that if this is really dangerous enough that we would want to pause, that incentive also applies to China. I feel like people talk about this as if China entirely consists of people unable to understand that, or intrinsically evil so that they just want to do whatever will be bad or something. They presumably also don’t want to be destroyed by AI or to have some sort of dystopia.

I do think that it’s more likely that someone in power there, like a human in power there, would want to have a human coup. Whereas here maybe power is more spread out, so that changes things somewhat. But yeah, I think we, in both cases, should be trying pretty hard to coordinate with them to not be doing this stuff.

I guess it just seems like trying to race them mostly increases the chance of all of this stuff happening, because it shortens timelines. I think we often talk about who ‘wins’ — and I think that’s sort of a bad way of describing it that we shouldn’t be doing.

Related episodes

About the show

The 80,000 Hours Podcast features unusually in-depth conversations about the world's most pressing problems and how you can use your career to solve them. We invite guests pursuing a wide range of career paths — from academics and activists to entrepreneurs and policymakers — to analyse the case for and against working on different issues and which approaches are best for solving them.

Get in touch with feedback or guest suggestions by emailing [email protected].

Our crash course on transformative AI

We've carefully selected 10 key episodes to help listeners get to grips with the potential upsides and downsides of powerful, transformative AI.

Check out 'The 80,000 Hours Podcast on AI'

Listen here, or anywhere you get podcasts:

If you're new, see the podcast homepage for ideas on where to start, or browse our full episode archive.