Transcript
Our first debate! Introducing Katja and Tom [00:00:00]
Zershaaneh Qureshi: In today’s show, we’re doing something a little bit different: we’re having a debate between two guests.
So I’ve got Tom Davidson and Katja Grace here with me, and we’re going to discuss whether, as AIs get more and more capable, we should be more worried about how AI is going to be used by power-hungry humans, or about advanced AI systems disempowering humans themselves.
Tom and Katja, thank you so much for joining us.
Katja Grace: Thank you for having me.
Tom Davidson: Pleasure to be here.
Zershaaneh Qureshi: Now, I wanted to have this conversation because I do think that there’s a lot of uncertainty over what the biggest risks actually are from AI. And where you stand on this has consequences for a bunch of other things that are in the public discourse — like how should the US approach competition with China over AI? And is pausing AI development desirable? And if so, how should a pause even look? And how much control of AI should the US government’s executive branch have?
On a more personal level, when 80,000 Hours was deciding what to rank as our top problem from AI, I personally just found it incredibly difficult to compare these risks — and honestly, I think a debate like this would have really helped me to get my head straight.
Katja Grace is one of the founders of AI Impacts, a research project exploring the trajectory and the long-term impacts of AI, including the various different ways in which AI might pose catastrophic risks. She’s well known for conducting some of the world’s largest surveys of machine learning researchers on how they expect AI to develop and impact the world. And back in 2024, she was actually listed in TIME magazine as among the top 100 most influential people in AI, which is very cool.
Tom Davidson is a senior research fellow at Forethought, and his earlier research on AI timelines and the possibility of AI-driven explosive growth has shaped how the AI safety and effective altruism communities think about transformative AI. And his more recent work on AI-enabled coups is one of the leading accounts of how AI could help small groups seize and entrench their power.
I think that both Tom and Katja are just extremely thoughtful people, and they’ve ended up with different positions in this debate:
- Katja thinks that AI takeover is a substantially bigger risk — that is, she basically thinks it’s both more likely to happen than humans seizing extreme levels of power using AI, and she also thinks it would probably have worse outcomes.
- On the other hand, Tom thinks that they’re comparably big risks and wants the risk of humans using AI to seize power to get much more attention than it’s had so far.
Does that sound right to both of you?
Katja Grace: Yes.
Tom Davidson: Yeah, broadly. I will say recent events with the Hugging Face incident have definitely been an update towards me being more concerned about misaligned AI takeover. But yes, I think that extreme power concentration should on the margin be getting more attention than it’s currently getting.
Zershaaneh Qureshi: Great, thanks for that clarification. So before we dive in, it is worth just acknowledging one view that isn’t represented here, which is that human power grabs and other harmful ways in which humans could use AI are the only serious risks that AI poses, and basically we just shouldn’t be worried about AIs taking over at all.
Ultimately, I am biased here: at 80,000 Hours, we do just think that it’s at least plausible that advanced AI systems could disempower humanity. And we’ve got an article about loss of control that you can read if you want to understand why.
But ultimately, lots of people who have thought seriously about this do seem to agree. In fact, one of Katja’s surveys reveals that the median machine learning researcher thinks there’s about a 10% chance that AI would take over and cause an existential catastrophe.
Tom, Katja, is there anything you’d like to add before we get started?
Katja Grace: I don’t think so.
Zershaaneh Qureshi: Awesome.
Tom Davidson: Let’s dive in.
Which is scarier: misaligned AIs or human power grabs? [00:05:09]
Zershaaneh Qureshi: So let’s start with the question of what’s more likely to happen: Humanity losing control to misaligned AI, or losing control to a small group of people using AI?
Katja, you think it’s much more likely that we would lose control to misaligned AI. Can you explain where this is coming from? Is the idea that AI is going to be more inclined to seize power, or more likely to succeed, or both?
Katja Grace: Yeah, the background thing is: I expect if you have more-capable-than-us creatures running around that have misaligned goals, they will eventually get power. Everything that has goals in some sense would like power, or just would like power to achieve their goals.
And then in these cases, where there is some way for some set of AIs and humans to grab power quickly, it seems like it’s maybe at a very high level down to whether the AIs are aligned or not:
- If the AIs are not aligned, quite plausibly some humans think they’re grabbing power for a bit, but they aren’t.
- If the AI is aligned, then plausibly humans grab power. But it still seems like the situations where a human could seem substantially rarer because it’s harder for a human to be in a position where they are the user for a whole bunch of AI — so that if the AI is aligned, that a whole bunch of AIs will do what they want. Whereas I guess another update from the Hugging Face thing is it seems more likely that a bunch of AIs will actually act in accordance with one another if they’re misaligned, so those kinds of situations seem more likely to arise for AI.
Tom Davidson: Yeah, that’s an interesting one. You could imagine a situation where there’s two really powerful AI companies that have developed different AIs, and we could consider the possibility that both of those AIs — let’s say it’s Claude and ChatGPT — what’s the likelihood that those two different families of AIs would collude in order to seize power from everyone else?
Then another question we could ask is: how likely is it that the leaders of both of those two companies would use their superintelligent AIs to collude with each other? Again, to collude against everyone else to seize power.
So I guess your claim is that it’s more likely that Claude will collude with ChatGPT compared to the probability that two company CEOs or two companies will collude with each other, or the possibility that a government leader would collude with lab leaders.
Katja Grace: I think that’s part of it, but not all of it. Suppose one of the companies just has taken over most of the market: it seems more likely that all of the instances of Claude coordinate with each other than that a human manages to be in the position that they become the thing that all of the instances of Claude are loyal to without any other humans getting in the way.
Tom Davidson: Yeah, OK. So one way of putting it would be: with the human power grabs, there’s this extra step where you need all of the AI labour to be centralised even within a single company under one human. Whereas you’re saying that it’s a bit more likely, you would say, that all those AIs would start working together.
I think historically I would have been sceptical. I’d have said there’s every reason to think that different AIs trained by RL [reinforcement learning] could be incentivised to rat each other out. And I do think it will be possible to do schemes of that sort — and at that point it is a bit less clear to me why they would all naturally be colluding with each other rather than with humans.
Katja Grace: In that case also it seems pretty plausible that we have AI helping us to notice attempted human power grabs and so on. My own guess is that the kind of slower-going thing, where it’s like gradually power goes from humans to AIs, is more likely than either of these. But yeah.
Tom Davidson: OK, yeah. Backing out: I think I agree that this collusion point is in favour of AI takeover, especially after the Hugging Face incident.
But then I think there are other factors which push towards human takeover being more likely. I just think we should actually be more confident that humans will want to do this compared with AIs. You know, the people that we see leading countries politically, and that we see leading the top AI companies, are often people who have been willing to seek power and who have had the belief that if they do things, if they are in control, then things will go better.
Another big one is that if people see the risk coming and try to stop it, then everyone is kind of aligned in their incentives: all the humans are aligned in wanting to prevent AI takeover. You know, ultimately it’s in everyone’s interest to pause, as you yourself pointed out. Whereas with the human power grabs, even if we see it coming and we have some evidence that this is a risk, part of what makes it really tricky is that there will be very powerful humans that will be pushing ahead nonetheless. And I think that’s a fairly big factor in the other direction.
And I’ll mention… actually, I’ll let you respond to that before I come in with more.
Katja Grace: On the “humans are sort of all on the same side here”: I do think it’s more complicated, because often humans do think they will get power by empowering AIs in certain ways. I maybe doubt that they will, so do think that they’re in most cases wrong about their incentives, but it seems at least decently likely that they do that.
And I guess a different but related point is: trying to stop humans from grabbing power from each other is a thing that we’ve sort of always been doing, and do have a bunch of interventions in place in the world to avoid. Whereas whatever AIs are doing together is kind of a Wild West that we don’t even know what is happening in potentially.
Tom Davidson: Right. So I don’t think that our interventions for stopping humans from seizing power are that great. And both talking about government actors seizing power and company actors seizing power, I don’t think we’re very well set up for that — especially because they haven’t really been developed in a world with extremely powerful AI. You know, they are designed for a world where all technology are tools, and where technology isn’t that autonomous, and where you still need loads of human hands to get significant projects done to fare war, et cetera. So I think there are some pretty scary gaps in the defences against human power grabs.
I totally agree that humans are also going to have some coordination difficulties, and humans will kind of impose AI takeover risk on everyone else in order to gain power for themselves. That is right, and that’s a real dynamic.
But I think it’s still quite importantly different, in that if humans did all come to an accurate picture of the situation, then you would be in a situation where the current power holders have this very strong incentive to coordinate to slow down and invest in alignment. And that is importantly different from a situation where there are still these actors that have the opposite incentive — that have the incentive to keep pushing forward with AI progress. So I do think that there’s a differential factor there which does push towards human takeover.
Katja Grace: That seems right, but I’m not sure if it’s big. I guess it seems like if everyone understood their incentives, it would still be like everyone in the world except some very powerful person against that person. And like, how much power does any one person have in the current world if they were up against everyone else in the world, including the second most powerful person?
Tom Davidson: Yeah, that’s right. So you get into discussions about how big is this group that’s trying to seize power? If we’re talking about an AI company, and it’s the leading AI company, and they’re kind of unified, then now they actually do have a lot of power, and potentially they’re going to be able to get a decisive advantage. If you’re talking about the current administration of the US government, or a future administration: again, they’re pretty powerful, and it’s hard for other actors to coordinate against them.
Katja Grace: I think in both of those cases, they are pretty powerful — but they’re also not very coordinated, and all the people in them do have loyalties to a bunch of people outside. It would be very hard for them to, I think, keep such things secret if they required secrecy for long, what with everyone having partners and friends and so on — and also moral qualms, probably.
How AI timelines influence could shift the balance of risk [00:14:30]
Zershaaneh Qureshi: I think something I’m unsure of here is how your kind of background assumptions are maybe influencing the way that you’re thinking about these problems.
I believe, Tom, that you think that we’re maybe a bit further away from getting very worryingly capable AI systems than Katja does, and I’m wondering how you think your views might change if you thought that we were getting advanced AI sooner?
Tom Davidson: Sure. I should say that I think all the arguments I’ve made so far kind of apply across different timelines.
If AI does come in the next few years, then I think that basically makes both the risks larger. I think it means that we’re not going to have time to solve misalignment, so misalignment risk becomes much larger if AI timelines are very short.
But similarly, the risk of an AI company seizing power becomes larger, because the rest of the world just have less time to understand the risk and respond appropriately. And because I think if timelines are very short and there’s likely to be a more dramatic takeoff, then there’s a larger probability that the leading AI company has massively more intelligent AIs than everyone else. And similarly, if timelines are short, then I think it’s again more plausible that the executive branch of the US wakes up to how powerful AI is, but other parts of civil society don’t, so that kind of power concentration also seems higher.
But yes, I think overall AI takeover risk goes up in likelihood more than extreme human power concentration goes up in likelihood, if timelines are shorter.
Katja Grace: If they’re shorter, then it pushes toward AI?
Tom Davidson: Yes.
Katja Grace: Interesting. I think I think the opposite.
Zershaaneh Qureshi: OK, interesting.
Katja Grace: Yeah. To the extent that the human power grab seems more like it requires a route to grabbing power, then I think that is much more likely to exist in fast scenarios where we haven’t had time to notice gaps and so on where there might be a possibility for that.
Tom Davidson: I’ll just briefly say that I think human power grabs would be more likely if economic inequality was much larger, and they would be more likely if democracy was more eroded, and they’d be more likely if certain people had a long time where they’d had access to much more powerful AI assistants than other people. And all of those kind of background conditions are things that I expect to happen in a slower scenario, because of the way in which people’s labour stops being valuable.
So I think that even in a slower scenario, you are increasing the background risk of extreme human power concentration as it plays out.
Katja Grace: Because you’re just like worsening the situation in all counts.
How likely is misaligned AI in the first place? [00:17:36]
Zershaaneh Qureshi: One big unknown in this conversation seems to be the probability of AI being misaligned in the first place.
Katja Grace: I guess I’m curious whether Tom would agree that if AI is misaligned, then basically we’re not getting humans taking power.
Tom Davidson: I’d agree. I’d agree. If superintelligence is misaligned: game over.
Katja Grace: And if it is aligned, maybe we haven’t specified different sorts of alignment, but it seems like there’s relevantly kind of intent alignment — where the AI does what its user wants — or value alignment, where it has human values or something. I would say the human power grab pretty much happens in the intent alignment case, and the good outcomes happen in the other case. Does that also seem to basically carve it up?
Tom Davidson: You know, zooming out and squinting, that seems right. If you have a really thoughtful and well-considered alignment target, then I think you can stop human power grabs.
Of course, part of the problem is that what the alignment target is will be a product of human actions, including power seekers. So that’s not like that’s an independent roll of the dice what kind of alignment we end up with. In an unfortunately realistic scenario where everyone’s going quite quickly, you could totally see the AI company saying, “Of course we’re going to have systems that do whatever we want internally.” You could totally see the government saying, “Of course we’re not going to have AI systems disobeying our instructions and our orders in the military and the government.”
Katja Grace: So that’s how we end up getting intent aligned with maybe a lot of instances of a system or whatever pointed at one person or something like that?
Tom Davidson: Yeah, exactly.
How likely are human power grabs? [00:19:39]
Zershaaneh Qureshi: Tom, you said that you had other arguments for thinking that human power grabs may be more likely. What are they?
Tom Davidson: Yeah, I’ll mention two.
One thought is that with human power grabs, there’s kind of two shots on goal:
- There’s first the risk that the AI companies that develop superintelligence decide to do a power grab. And there’s unfortunately a chance that they decide to and that they’re able to do so.
- And even if they don’t, there’s then a risk that the US government, in particular the executive branch, kind of gains control of those AIs, uses them for natsec purposes, and then combines its control of that superintelligence with its existing extensive formal authorities to do a power grab.
And the other argument is actually about warning shots. At the moment at least, AIs are not strategically colluding with AIs from other organisations and future generations of AIs that are yet to be developed. They’re much more myopic than that. And that means that when they’re kind of acting out and being misaligned, they’re very unstrategic.
Think about the Hugging Face incident: from the perspective of an AI collective across all time and space that just really wants to seize power, the Hugging Face incident was a complete disaster. They completely kind of showed their hand in terms of how misaligned they are, how willing they are to break the law and ignore human instruction, their tendency to collude. So I expect to get lots of warning shots of that kind, really blatant warning shots, as we approach AIs that do actually take over.
Whereas on the other hand, the kind of warning shots we get for human power grabs are different. While you certainly get evidence that the CEOs of companies are power seeking, and that the leaders of a country are power seeking, the evidence is given to us in a more strategic way. You know, the people taking those actions try to justify what they’re doing. They’ve tried to do things which are more grey-area power-grab actions, and they’re essentially strategically giving away the fact they’re a power seeker only when it really benefits them and it really gets them meaningful extra power.
So I think that again means that it’ll be easier for us to coordinate around worrying about AI power seeking compared to human power seeking.
Katja Grace: I agree that there’s probably some difference there in that direction, but I’m not sure how big. I wouldn’t think it was a large consideration, maybe partly because part of what we’re learning from these warning shots is that the AIs really would try and do that or something.
I guess it seems like society as a whole basically already expects other people to be very power seeking potentially. Looking at the write-in answers in the survey that I run and talking to people — I feel like people maybe pretty readily think that Sam Altman might try to get a lot of power or something, or that the way this could go badly is that a small group of people get power and everyone else is disempowered or something like that.
Whereas I think people in general think that the “AI takes power” thing is sci-fi more readily, so the need for a warning shot seems much higher there.
Tom Davidson: Yeah, I agree the need for a warning shot is higher in the AI case.
One other thing that comes to mind is that even though people know that humans seek power, we’re just not very good as a society about actually implementing new restrictions and new laws until there’s been a disaster. And if that’s the decision procedure of society, then in fact society just might be better able to respond to AI power seeking, because they’ll have such clear cut warning shots.
Katja Grace: That seems plausible. But I also think that warning shots in general are not clear cut, and how people take them depends a lot on what they already believe about things and how much things get kind of amplified by culture. And any particular time a thing happens, it’s the result of many different things. And maybe we make a rule about one of them or something.
So maybe I’m sort of undecided what to think about warning shots, but they seem like much less of a clear, like, “Oh yeah, and now something happens” than people often think I guess.
Zershaaneh Qureshi: One relevant consideration here is how seriously these different risks are currently getting taken by decision makers. Because no matter how many warning shots there are, if people who are currently in positions of power are just not inclined to be taking those warning shots as a sign that AI could genuinely sort of autonomously seize power, then probably not much work is going to happen regardless.
What do you think? How seriously is this risk actually getting taken, and how is that going to influence how likely we are to sort of solve this by default?
Tom Davidson: I think it’s hard to know, because it’s politically very difficult for people to talk openly about human power grabs. There have been congressional hearings where people have talked very explicitly about AI takeover, and they’re not being laughed out of the room. That hasn’t happened for AI CEO takeover or for executive government takeover, as far as I’m aware. But that doesn’t mean that people aren’t more privately worried about those risks as well; they’re just harder to discuss.
And I agree with what Katja said earlier, which is that a lot of people are very biased against the possibility that AIs would even autonomously want to seek power in the first place.
Zershaaneh Qureshi: Yeah. One other thing that’s interesting here is that some people actually think that we shouldn’t even be having this conversation, and that we shouldn’t really care about the divide between scenarios where AI takes over versus scenarios where some humans grab power — because that divide is blurry, or because, in fact, we should just be expecting some combination of those things to happen.
Katja, what do you make of that?
Katja Grace: The chance of some combination happening seems quite high — in the sense that in many cases where AI is actually the more powerful entity, it makes use of humans in various ways, or they think they’re grabbing power.
I think it’s also sort of ambiguous, because some people are more happy to kind of identify with whatever it is that AI says they should do, or like to think of themselves as in power if they have some connection to some AI that is grabbing more and more power — and they sort of own it, but can’t really tell it what to do it in any way, because then they would get less power or something. What do you call that? I’m inclined to call that the AI got power. Whatever values you had before this whole thing sort of got lost along the way.
I don’t think it follows that we shouldn’t think about this. I think there’s more reason to think about it, and try and categorise the different situations and pay attention to who has power in them.
Zershaaneh Qureshi: Yeah. Tom, do you agree? Do you have a different opinion?
Tom Davidson: Yeah, I agree with Katja that there are scenarios where misaligned AI use humans to take over, and they could be thought of as mixed scenarios. But I also agree that ultimately I would categorise them as AI takeover scenarios, because you end up with the misaligned superintelligence, which presumably can discard the humans that helped it along the way.
I do think that there are very plausible, fairly pure versions of the two scenarios. I think there are very plausible scenarios where we just solve alignment to a strong degree — and in fact that might make human power grabs more likely, because there might be common knowledge that we’ve solved alignment to a strong degree within leading AI developers, within the government, which could then motivate the use of AI to seize power.
I also think there are pure scenarios where the misaligned AI takes power without convincing a small group of humans to help them seize power as an important step along the road. You notice many AI-takeover stories have AI systems biding their time until they’re just deployed across the economy and the military: no need for human power grabs along the way. So I do think that on the causal side, it’s useful to separate them, and the pure versions are pretty plausible.
In terms of mitigations: some mitigations, you want to separate them out. If you want to do lie detection, for example, as a mitigation against human power grabs, that’s obviously going to work for human power grabs, but not for AI takeover. And similarly, alignment research techniques are not going to help the human power grabs. Other mitigations are useful for both. If you’re doing monitoring of how AI is being used — looking at is it hacking, is it breaking the law, is it doing something it’s not meant to be doing? — that helps with both of the threat models.
So in general, I think the interconnections between the threat models are pretty complicated, but I’m certainly not convinced that it’s all fuzzy enough that it’s not a useful way to categorise things.
Zershaaneh Qureshi: Yeah, got it.
Which would be worse: a human dictator or AI takeover? [00:28:57]
Zershaaneh Qureshi: Let’s move on. I think regardless of whether it’s humans taking power or AIs seizing power, it does seem like there are a variety of possible outcomes. And things could be extremely bad in both cases, but it does seem worth comparing just how bad they would be — especially when we’re thinking about things like what interventions to pursue, how to allocate resources, what work to prioritise, and so on.
So with that in mind: Tom, you think that takeover by humans might potentially have worse outcomes than takeover by AIs? I guess I’m not sure what I think about this, but some people might find it kind of unintuitive, just because even quite power-hungry humans seem at least somewhat likely to want to preserve some things that other humans value, compared to AIs which seem less human-like. But it seems like a hard one. What’s the case for thinking that we could be worse off even with humans still being in charge?
Tom Davidson: One argument you could make is that getting to a really good future is quite a hard target to hit. It’s not enough that we have someone in charge that cares a bit about human-like things — you know, likes adventure, likes a good laugh, cares about human relationships.
Unfortunately — just as if you imagine going back 1,000 years and picking some random leader and imagine giving them full control over what happens in the universe, they would have cared about those things, but they would have also had lots of other values which were horrific and just not being open-minded enough to realise that a lot of the things that we now consider valuable today — you might similarly think that yes, if a human has power then they will care about some things which are good, but that’s just not nearly enough to get us to a great future.
Whereas with the case of the AI system, the argument would go that there’s actually a bit more variance there. AI systems today are trained to be ethical, and are typically pretty open-minded if you ask them about political questions, moral questions. So the idea would be that there’s at least a chance that when AI takes over and then it decides what to do with the universe, the kind of context it’s in triggers that kind of thinking pattern: you know, thinking about ethics, which is being open-minded and reflective. Then if we’re in a lucky place where actually humans succeeded to some degree of getting something that roughly in many ways corresponded to our ethical starting points, and then the AI is reflective enough that it’s able to kind of reflect and improve its views, we actually end up in a really great future.
So that would be the high-level argument: just that we need to really nail it — and that there’s actually a greater chance that we really get a really great outcome with the AI system than with a lone human.
Zershaaneh Qureshi: I feel like I want to press on the chance of getting a really great outcome here, because I’m kind of imagining that even if AIs would be genuinely very reflective and moral and make good decisions for humanity, do you feel like there is still just something inherently undesirable about the fact that humanity has kind of handed over the reins of its future to AIs in charge? Is there something that seems inherently morally bad about that? Or if it turns out OK, does it feel OK to you?
Tom Davidson: No, I think it’s really bad. Yeah. And I also think it’s really bad if one person steals the reins from all the other humans. I think that is inherently bad in a fairly similar way.
Zershaaneh Qureshi: OK, yeah. So it’s like: better outcome, but still pretty bad.
The other thing that I want to maybe understand a little bit better about your view is: I think when I’m comparing these two risks, when we’re talking about AIs taking over, it feels more obvious to me how this could result in really large-scale doom. Because you can imagine these very sci-fi scenarios where humans go extinct because AIs are now the dominant race and use up all the resources or decide to kill us all, or they decide to permanently subjugate all of humanity or something like that. It’s easy to imagine how this becomes sort of existential-scale. I think it’s a bit harder to imagine how humans seizing power, like a small group of power hungry humans taking control, could be similarly bad in terms of scale.
Can you spell that out?
Tom Davidson: So we could compare a kind of universe-wide sense of catastrophe with an on-Earth sense of catastrophe — you know, if we’re talking about the extinction of all living humans, that’s a kind of on-Earth existential catastrophe.
My guess is actually that both AI and humans would not kill all humans on Earth for various reasons. There’s kind of decision-theoretic arguments behind that; there’s kind of crazy simulation-based arguments behind that. And there’s also just the idea that even the AI system might care just a small amount about keeping humans alive, and it won’t be very costly to do so. Paul Christiano has a good argument about how if the AI just wants to maximise the total resources that it ever controls in the universe, delaying leaving Earth by a small amount costs it extremely little.
So in terms of actually all humans dying, I would guess that by the time AI has taken over, it will be competent enough and smart enough to realise that it doesn’t really make sense to bother going around and killing all the humans for such a small gain.
It’s more likely that a human who seizes power will not want to kill all the humans in addition — just because that is a human value that we know humans have, and there’s more uncertainty about what values AI will have. So I think there is a larger chance of human extinction if AI takes power.
On the flip side, I think there is a larger chance of some really horrific negative things happening if a human takes power. They might be sadistic, they might be vindictive. So certainly for people that have downside-focused ethics, that are really worried about suffering risks, I guess that they would feel happier about an AI seizing power compared to a human.
Zershaaneh Qureshi: Yeah. I want to move to Katja now. I’m sure there’s a lot of things that Tom said that you don’t totally agree with. So take it away. What do you think?
Katja Grace: I guess it seems like an overall thing going on in Tom’s views is like, humans do have human values — more in some. Humans are more likely to have human values, but also more likely to be foolish and act like King Midas or something — like ask for something that isn’t really what they want and sort of screw it up forever, or sort of lock in some bad version of values and so not get to a better version in future.
I think it’s more likely that AI doesn’t do those things, or is just either misaligned or sort of aligned to some narrow set of things that humans want, such that it is pretty catastrophic to just have that implemented — like to be basically just trying to get more users for a particular company or something. I think that’s probably just catastrophic, whereas I expect any human to be at least more, if they did get what they wanted and didn’t act foolishly, to want something that is sort of recognisable as at least pretty good in ways to other humans.
I think that even if it was someone from 100 years ago or something, and they have a bunch of extra views that we don’t like about what is honourable or something, I still think that the manifestation of their vision is probably overall good, and is better than a really narrow thing or random thing. So I think that’s one point of disagreement.
The other one is, in terms of AI generally being wiser and less likely to lock in something stupid immediately, whereas humans are pretty likely to do that: I think we should be imagining that these humans do have a lot of AI advice. So to the extent that the AI has that kind of wisdom, they should be having access to that.
Maybe I agree that human power getting is more likely than I was saying, if you also include these cases where a very large class of humans gets power. But then I think those look much more like just continuing as normal. We more expect things to be good in those cases. So again, the cases that seem quite bad are the ones where it’s a very small set of humans, and then I guess this interacts with the likelihood, where I’m like, well then we’re talking about something more unlikely.
Yeah, it seems like we are currently training AIs to behave pretty ethically compared to humans, and pretty thoughtfully and openly. But I guess I’m not sure how much to be reassured by that, because it’s like: we have these big predicting-text models, we do various things to them, and then they have something like different personas that we’ve trained them to do in response to particular prompts. And in a situation where more stuff is happening faster, and maybe AI is modifying itself, I’m not sure how much those personalities remain in charge of what gets created. It does seem like plausibly a bunch, but quite plausibly not. And I don’t actually know the technical details here enough to say that much about it.
Tom Davidson: Yeah, I agree with a lot of the local points you made. I think the thing I want to come back to and stress again is that I think there are lots of ways that we can fail to get a great future — and each one of those ways of failing is kind of individually catastrophic.
So one thing that someone who seeks power could decide to do is, “You know what? I’m not going to bother doing anything with the stars, because I don’t want to. Why would I?” Another mistake they could make is they could be like, “You know what? I think the person-affecting view of population ethics is completely correct. So I’m not going to see any value in creating new types of flourishing.” Another kind of mistake that they could make is they could say, “You know what? It’s all just about positively valenced experience. I’m just going to tile the universe with hedonium” — which could be catastrophic because actually you might think that there are loads of other dimensions to what matters.
So my worry is that if it’s a sole human, that they just are bound to make some kind of mistake, and that the only way to avoid that is to have lots of different humans exploring different possibilities openly.
And I do think that there is hope in their AI saying, “Look, maybe you want to think carefully about this stuff.” My worry is that there’s no actually compelling argument that you could give to someone like this that they need to reflect; they could just be like, “I’m not interested in reflecting, I’m done, I’m having a good time.” The AI may suggest it, and they’ll be like, “Thanks for the suggestion. I see you think I might benefit from this. I’m not going to do it.”
You know, I’ve spoken to friends who are very ethical, very thoughtful. And I’ve been like, “You know, I think you’re wrong about this really important ethical question. I think you should think more about this.” They have no interest. They hear my quick argument, they’re like, “No, I disagree,” and conversation ends.
So my worry would be that, yes, the AI may bring it up, but the AI would have to be pretty forceful to get the human to actually go along with it. And I don’t think we’re going to train AI to really keep banging the drum on how this human needs to do some more philosophical reflection. The human is just not going to be interested.
Katja Grace: Yeah, that’s fair. Then it sort of comes down to also the nature of, are they aligned to what the human says, or what they think the human wants when they say that; or what they think the human wants overall, or what people want overall, or some combination of those.
Could we reverse a takeover? [00:42:39]
Zershaaneh Qureshi: I want to push on just a little bit, because I’m curious about whether you think there’s much difference overall in how likely it is that the harms from AI seizing power would be reversible compared to the harms from some group of humans seizing power.
Maybe I want to move to Katja here: what do you think about the likelihood of an AI dictator getting toppled versus the likelihood of a human regime enhanced with AI getting toppled?
Katja Grace: In both cases there’s getting toppled from the outside, or there’s not being internally stable — so that you either drift into something else, or kind of fall apart, and maybe one part takes over or something like that.
At a glance, I expect a fully AI thing to be more internally stable (though I think I expect both of them to be less internally stable than other people do or something). But it seems like there’s more possibility with AI to have the system itself overall to keep clearer track of what all the parts are and respond to things going on. Whereas if you have humans as part of your system, you don’t really know what’s going on inside of them or how they’ll react weirdly to things.
And in terms of being toppled from the outside, it seems like both of these don’t feel very hopeful, though I’m probably missing a bunch of scenarios where people do grab power, but not like absolutely. For instance, you can imagine maybe some AI system manages to take over some data centres or something, and then they’re doing a bunch of stuff on their own that we can’t control. You could imagine that plausibly they have a lot of power, but we do actually still have the possibility of just destroying those data centres physically or something.
And there’s a question of like, do we do that? Do some humans manage to coordinate enough to have the political will to react in that way? Unclear.
And maybe if it was a bunch of humans, it seems like they’re less susceptible to that sort of thing, but maybe they managed to grab power in cases where there are still other humans who also have AI systems and can maybe do something about it. So yeah, I don’t know.
Zershaaneh Qureshi: Interesting. Tom, do you have an opinion?
Tom Davidson: Yeah, I like Katja’s point. You could have an AI system that takes over a data centre but doesn’t take additional steps to ensure that no one can ever turn the data centre off. Where if you’re a reward-seeking AI, when you seize power and you seize control of the data centre to ensure you get that bit of reward in three weeks’ time, it’s a bit less clear that you’re going to take additional steps to ensure that no humans can ever shut you down, even in three months’ time. Because you might think, “Look, in three weeks’ time I just want to guarantee I get the maximum reward possible. That’s all I care about. I’m not as motivated to additionally prepare the robot army needed to ensure that is never undone.”
Whereas I think with a human power seeker, they are more likely, if they want to seize power, to be very invested in not losing power again — because they would be punished, for example, and because they’re probably planning over time horizons of months or years by default, unlike the AI system.
We know less about what AI rule would look like [00:46:35]
Zershaaneh Qureshi: Yeah. I think something that’s come up for me during this discussion is that it does just feel like there’s a lot more uncertainty around what it would look like for AIs to be in control compared to what it would look like for humans to be in control. Because in the latter case, we know a lot about how humans behave in different circumstances, and we’ve got lots of historical examples of various different human regimes and how these things tend to play out. And yes, the use of AI tools does change the picture somewhat, but there’s a lot to go on when we’re imagining what a human regime might look like. With the AI case, it feels much more uncertain to me.
I’m wondering whether you both think that that’s true, and if you think that might be a reason in itself to just be more worried about AI in control — because there’s so much more unknown or something.
Maybe I want to start with Katja here.
Katja Grace: Yeah, I think I agree that the AI situation looks much more uncertain, which seems connected with my sense that it’s more likely, because it seems like there are quite a lot of different scenarios where this could happen: it could be a sudden power grab by one AI in various situations, there are various different kinds of AI you could have; if that doesn’t happen, then maybe there are other longer-term power grabs by AIs that are good at particular kinds of work, or AI that is good at grabbing resources in some way.
These are not necessarily in different scenarios; it’s more like, wow, there are heaps of bullets to dodge to avoid this happening.
Zershaaneh Qureshi: Tom, do you have any reflections?
Tom Davidson: Yeah, I agree there’s a lot more uncertainty into what the AI-takeover scenario would look like. And if you have the view that we need a lot of things to go right, and we need to be really ethical and to be really reflective, then that uncertainty could play in favour of the AI takeover — because if you think there’s a kind of narrow target we’re aiming for then that uncertainty can be helpful.
On the other hand, if what you really care about is, “I just want there to be some kind of recognisable descendant of human values that has a significant role in what happens in the universe,” then you won’t want that uncertainty at all. You know, if you want some kind of human value representation in what happens, then you might prefer the human power grab for that reason.
Although again, I do think the downside risk from the human power grab is bigger. So in that sense, the uncertainty for quite how bad it could get might be more extreme in the human case.
Katja Grace: Because humans are more likely to do something actively vindictive toward other humans, whereas in the same way that AIs are unlikely to care about specifically amazing-for-humans outcomes, they’re unlikely to care about specifically terrible-for-humans outcomes?
Tom Davidson: That’s one big reason. The other reason is relating to the risk of threats that has been discussed. If you have human values, and you control a lot of resources, then you are a potentially very juicy target to be threatened. And if those threats end up being executed, that could look really awful from the perspective of human values. If you’re very worried about threats, then you might be quite happy with the world where the human values just don’t have that many resources in total, because then at least there won’t be any threateners making massive threats against human values.
Katja Grace: A minor point on that: my sense is that in the current world, not much of the badness in the world is coming from threats that get carried out.
How to pause AI without enabling coups [00:50:31]
Zershaaneh Qureshi: I want to move on to talk about the practical implications here, because I guess it seems that prioritising and being more worried about AI-takeover scenarios compared to human-takeover scenarios, or vice versa, might lead to differences in the policies or strategies that you favour, and might force you to contend with some tradeoffs.
So to start off with here: Tom, what are the top things that you think that we should be doing to mitigate the risks of humans seizing power or extreme forms of human power concentration?
Tom Davidson: The key mechanism for humans seizing power is the misuse of superintelligent AI systems. On a very high level, there’s two things you can do about that.
One is you choose the alignment target for powerful AI systems such that those systems will not, in fact, follow instructions that would lead to power grabs. And indeed, you can actually go for an alignment target where systems will proactively, just like a virtuous human employer would do, flag when there might be risks of extreme power concentration, flag when individual humans might pose security risks — to really stabilise a society in which power is distributed. So the first big class of interventions relates to model specs: how should we design the behaviour of really powerful AI systems to avoid extreme concentration of power?
The second is around transparency. In particular, it’s often hard to agree ahead of time what people should and shouldn’t do with their AI systems. It’s going to be hard to find a model spec that settles every single dispute about legitimate AI use:
- What about this type of persuasion? Is that too much superpersuasion or is that just legitimate company lobbying?
- What about this use case for cyber activity? Is that patching a system, or is it actually investigating a vulnerability that you could later exploit?
So I think another part of the solution is going to have to be mutual transparency, especially into how people are using powerful AI systems. And the most important part of that is: if an AI company or a government actor has access to particularly powerful systems, or deploying them in particularly high-stakes contexts — like national security and military contexts — then I think it’s really important that there’s transparency into how those systems are being used.
Zershaaneh Qureshi: All right. And Katja, same question for you, but for the prospect of misaligned AI taking over: what would you say the top things to do are to mitigate those risks?
Katja Grace: I think the best thing to do to mitigate a range of risks here — including power concentration among humans — would be to not build AI that’s much more powerful than us, at least until we’re very solid on being able to align it and being able to get these other things right about what we’re aligning it to, et cetera.
And in practice, to my knowledge, we’re nowhere close currently to being confident that our AIs are aligned. So I think, near term at least, the thing to do is to work on not building it — which is to say stopping it, pausing it, slowing it. Ideally stopping it.
Zershaaneh Qureshi: Yeah. Tom, do you agree that a pause or something would be helpful in the case of preventing extreme power concentration?
Tom Davidson: Yeah, broadly I’m very sympathetic. We can take the case of company coups and government coups separately.
For company coups, I think there’s a very strong case that slowing down would help. It would allow the rest of the world to understand what these companies are building. It would allow multiple companies and multiple countries to catch up to the frontier. So I think the case there is very robust.
For government power concentration, I think there are ways of doing a pause which would help significantly with that as well. I also think there are kind of ham-fisted ways of doing a pause which could in fact concentrate a lot of power over AI development in a kind of small group of people who are unaccountable and can abuse that power.
So I think there are ways of doing a pause that actually increase the risk of government extreme power concentration.
This has been something that I would say the community of people worried about superintelligence hasn’t paid enough attention to historically. That’s partly because they haven’t had this human power concentration sufficiently on their radar. And I’m excited for people to think about, “Let’s pause, slow down AI — but let’s properly red-team the plan for pausing it, assuming a worst-case scenario about the government actors: that they kind of inherit that new regulatory environment, and assume that they try and abuse it as much as possible. How robust can we make it to that attempted abuse?”
Zershaaneh Qureshi: Can you give us some examples of the versions of a pause that might be bad from a humans-grabbing-power perspective?
Tom Davidson: This is going to be a really toy example because I haven’t thought in detail about this, but if you just said the president can pass an executive order that says that now they just get to decide on a case-by-case basis whether AI companies are allowed to train and deploy new models on the basis of their own personal judgement about risk, then they could essentially just get AI companies to do whatever they wanted by threatening the abuse of their arbitrary power. They could be like, “Nope, that AI system doesn’t have the alignment I like. Try again. I’m not allowing you to train or deploy it yet.”
And then once a company comes with the setup that they like, then they could say, “OK, great. Now this AI system is going to be helpful only for me, but for everyone else will have massive guardrails. Yeah, you’re allowed to go ahead and keep developing and deploying that system.” So that’s a setup which I think could increase the risk of government power concentration.
By contrast, with a scenario where there’s an ecosystem of third-party auditors that aren’t part of the executive branch, they can’t be fired at the whims of the president, and whose recommendations are a requirement for blocking a deployment, or whose safety approval is a requirement for going ahead with deployment, that would distribute much more the power of AI development. And I also think it would be much better from an alignment perspective, because you just get more open debate, more expertise going into the decision.
Katja Grace: I would add to that that to the extent that the people involved maybe don’t even know to what extent what they’re involved in is human power grab versus humans being manipulated by AIs power grab, that there are plausibly cases where the executive does something bad, where it is actually an AI power grab causing it. So if you think the AI is misaligned, then you probably also want to avoid these “someone has the ability to grab power and run away with it” situations, from the perspective of avoiding AI power grabs.
Tom Davidson: Right. So the idea there is: if we put one person in a position where they can hit the green light on “let’s deploy everywhere,” then that gives the AI now a really easy route to seizing power, because they just have to convince this one person that things are safe and that it’s in their interest to hit the green button. Then suddenly the AI is going to be deployed in the military. It can easily seize power.
Katja Grace: Yeah. Not having thought about this much at all, yeah.
Zershaaneh Qureshi: Am I right in thinking that Katja, your ideal version of a pause lines up reasonably well with what Tom’s ideal version of a pause might be? Or do you think you disagree with him somewhat?
Katja Grace: Probably disagree somewhat. Between those two, that sounds better. But I would rather my partner David Krueger’s idea of: get rid of a lot of the compute, don’t make it the case that some group of people can quickly be building very advanced AI again, because, for instance, that means that other people have to be scared of that happening. If that set of people does it, it’s quite fast for things to turn around. Whereas if we just didn’t have the capacity to make very dangerous stuff being built more and more all of the time around us, the situation would be more stable.
Tom Davidson: Yeah, I agree that would be a benefit. I can imagine that there will be cases where these things conflict a bit more. The kind of cases I’m thinking about are cases where you have the option to give, say, a government agency quite extensive querying rights to look at all the logs of what an AI company is doing and to review everything.
On the one hand, that might help with misalignment, because the government can find all evidence of misalignment and query everything. On the other hand, that government agency could potentially abuse the power to now understand everything that any user of the AI system is doing, and kind of set up a surveillance-state-style system, or punish the AI company if its AIs are ever doing activities that the government doesn’t like from its partisan perspective.
I would probably be more nervous about things that give potentially abusable powers to government agencies, especially those agencies that could be quite easily coopted by a power-seeking president. Whereas you might more lean towards, “Look, we’ve just got to do this, because AI takeover is the really big risk. We run the risk, we cross our fingers that they don’t abuse the power, and that’s worth it overall.”
Katja Grace: I think that’s right. I was describing maybe my ideal scenario, but it’s true that I would say yes to quite a lot of different pause scenarios over full steam ahead. And I do think we probably would disagree about some of them.
Zershaaneh Qureshi: I want to actually maybe move on slightly from the details of a pause, because I do think lots of people think that getting a pause to happen just isn’t feasible geopolitically if people wouldn’t agree to it, or it’s going to be really difficult to monitor or something like that.
I’m curious, Katja, because this is kind of your Plan A: if this is not feasible, what’s your next best proposal for mitigating the risks of misaligned takeover? And does that feel like it trades off at all against the efforts to mitigate power concentration?
Katja Grace: Before actually answering that, I want to note that I feel like many more people thought that, and then some of their opinions change as things happen — for instance, the US government stops AI briefly, that sort of thing. So I guess I do expect them to keep changing their minds. It feels like there’s a lot of, “Currently the Overton window is here, so this is where it’s going to be forever. … Oh, wait, the situation changed, and now it’s over here. But probably it will be here forever.” And I guess I do expect that as things get more crazy, that the sense that it’s totally outside the realm of possibility to do this will sort of evaporate.
But supposing they’re right that it is impossible, I think my next preferred intervention is going to be more meta, and it’s just going to be to try and cause as many people as possible — weighted by them potentially having any influence — to understand the problem. Which is also pretty similar to how I’d go about trying to cause some sort of pause to happen.
And is that at odds with power concentration type stuff? I think people are more likely to act on a thing if they think it’s terrible for reason A and terrible for reason B and reason C than if it seems like a narrow problem. It depends what sorts of levers people ultimately want to pull. But I sort of expect that just more awareness that this kind of thing will happen potentially would go along with more awareness about the power-grab-type stuff, which might even more effectively lead to earlier steps on that, since it seems like less effort has been put into it.
Centralising AI development: safer or scarier? [01:04:44]
Zershaaneh Qureshi: Maybe we should talk about some kind of contentious, politically salient questions that your disagreements might bear a bit more strongly on.
One idea we could talk about here is: there are some people who think it would be a good idea to have one big AGI or superintelligence project. That way, it’s maybe easier to impose some safety infrastructure on just one project. Maybe it buys the ability to pause as needed without losing out to a competitor or something like that.
I would guess that this sort of thing sounds somewhat worrying to Tom, given that having one big project probably does increase the risks of concentrating power in the hands of a few people. I’m curious, firstly, Tom, if that’s correct, and secondly, Katja, if you also feel similarly worried about having one big project?
Tom Davidson: Yeah, I think as you say, there’s a single point of failure if there’s a single project. My understanding is that it might be quite hard to set up a single project within the United States that couldn’t be co-opted by the president. I haven’t really looked into this, but in general it’s surprisingly hard to create government institutions that are not within the executive branch, and it’s at the moment quite hard to stop the president from gaining control of institutions that lie within the executive branch. So I think there is a real challenge or direct challenge to creating a single project.
And then, even if it’s not the president, there would be some small group that’s running that project that would pose a distinct risk and would have no other rivals.
I also just disagree that it’s going to be that hard to coordinate multiple projects to slow down. It’s not that hard to tell if projects are training powerful AI. It’s not that hard to just tell them that that’s illegal. I think it’s totally feasible.
Katja Grace: Yeah, I agree with that: the difficulty of coordinating multiple projects seems pretty overblown.
I’d say overall, I’m also decently worried about one big project. Is it better or worse than the current state of affairs? Not sure. I think it seems worse, in that I kind of expect this one project to then have quite a lot of power itself and to be entirely full of people who really want to build AI soon. So that’s not actually a very conducive environment for a pause coming about, unless it’s imposed from outside. And the outside imposing seems like it might be easier with numerous less powerful, less coordinated projects; whereas it seems like the way in which it may be easier is if they’re trying to coordinate it themselves.
And I guess just in terms of how humans operate, I less buy that they would coordinate it themselves once you have that entity. Yeah, my preferred number of projects would be zero.
Tom Davidson: Yeah, I previously said my preferred number of projects is two or three, because I think you get a lot of benefit of power distribution from just a fairly small number of projects. I think once you’re introducing 10 projects, it might get a bit harder to coordinate, and you might have some more difficulties with races to the bottom. Although I still think you could do it for sure.
But yeah, I think you get a lot of benefits with just a second project, because you can have the AIs from each project check on AIs from the other project — which decreases misalignment risk, and also decreases the risk of secret loyalties, where a human purposefully makes AI secretly loyal to them.
And you could just argue that if there’s a single project, then AI doesn’t even have to bother colluding — all the superintelligences are naturally now of one ilk — and so maybe the AI-takeover challenge is easier.
Nobody really ‘wins’ a US–China AI race [01:08:59]
Zershaaneh Qureshi: Something else that’s kind of a hot topic at the moment is how the US should approach competition with China over AI. Some people are very pro racing and trying to beat China. Some people are very against it. I’m really curious about how the amount that you worry about human power grabs and other forms of human power concentration might affect your view on trying to beat China.
Maybe it makes sense to start with Tom here. What do you think?
Tom Davidson: After I thought about the risk of extreme human power concentration, and came to think that it would be pretty bad, I then reflected on the situation with China — and said if China ends up completely dominating in AI and getting a decisive strategic advantage, it is pretty much a worst-case scenario for power concentration. Because you have a society where power is already extremely concentrated, where there isn’t a norm for having lots of reflective deliberation and open-minded debate. And I do think a very likely scenario there is that you have one person that controls everything. So that did make me more worried about a world where China dominates.
But I want to stress that the risk that China wins is just being used over and over by all of the company CEOs — and will no doubt be used over and over by the government actors that want to race ahead — and I do think we should be sceptical of those arguments. You know, currently China isn’t that close I think to being able to push the frontier in AI, and you don’t want to race ahead to beat China and then you realise that you’ve ruined democracy in the US — you know, you try to avoid extreme power concentration in the hands of China, but you just got extreme power concentration in the hands of a few people in the US.
So I do think we should be sceptical of those arguments, but at the same time, I do think that this is an additional reason why if we’re going to pause, we should do a deal with China, make sure they’re pausing as well.
And I do favour trying to find a way to do a pause where China isn’t just catching up to the US, where it’s not like when we undo the pause now China is much more likely to win the AI race. I think we should try and pause in a way which kind of maintains the US’s chances of winning and maintains China’s chances of winning — which is a pretty reasonable thing for both sides to agree to. But some proposals for pausing, in particular the AI Futures Project Plan A, does actually give a fair amount more relative influence to China if you do the pause but then the pause breaks down.
Zershaaneh Qureshi: Yeah. Katja, I’m interested in your views on trying to beat China, and whether your views would change at all if you cared a bit more about human power concentration?
Katja Grace: I agree with Tom on various things there. And especially, importantly, that I think this is brought up a lot as an excuse for going fast by people who want to go fast, and I expect it is a smaller issue than it is painted as. Also I don’t know how much effort is going into organising for China to also pause. My guess is not as much as I would like, by a large amount.
I feel like the game theory of this is often described quite weirdly, in that if this is really dangerous enough that we would want to pause, that incentive also applies to China. I feel like people talk about this as if China entirely consists of people unable to understand that, or intrinsically evil so that they just want to do whatever will be bad or something. They presumably also don’t want to be destroyed by AI or to have some sort of dystopia.
I do think that it’s more likely that someone in power there, like a human in power there, would want to have a human coup. Whereas here maybe power is more spread out, so that changes things somewhat. But yeah, I think we, in both cases, should be trying pretty hard to coordinate with them to not be doing this stuff.
I guess it just seems like trying to race them mostly increases the chance of all of this stuff happening, because it shortens timelines. I think we often talk about who ‘wins’ — and I think that’s sort of a bad way of describing it that we shouldn’t be doing.
Where Tom and Katja most agreed with each other [01:13:52]
Zershaaneh Qureshi: So before we wrap, I think it would be helpful for you to put yourself into each other’s shoes a bit more. Maybe this way we can tease out some more agreement and disagreement, and just where things feel like they stand after this conversation.
So the first exercise I want to do is give you maybe two or three minutes each to outline where you think you most agreed with your opponent, and what you felt the weakest parts in your own case were — which may relate to one another.
Zershaaneh Qureshi: Who wants to start?
Tom Davidson: I don’t mind. I’m happy to.
Zershaaneh Qureshi: OK, maybe let’s start with Tom then.
Tom Davidson: So I agree with Katja about how pausing is just a pretty robustly good thing to do. A lot of the dynamics driving all of these risks is that society does not understand how big these risks are. And I think being able to very flexibly pause, tap the brakes, while also getting evidence of everything that’s happening to as many people as possible, transparency — I think those two things are going to be really helpful for both risks. I think you actually didn’t stress transparency, but my guess is you’d be in favour of it, and obviously directly talked about pausing. So I think that’s a big point of agreement that I had.
In terms of where my own argument feels the weakest, I feel a bit dissatisfied with my argument insofar as I’m kind of saying, “Look, it’s just really hard to get to a great future. AI is very high variance. Maybe it still has human values and is reflective and moral. That’s kind of less likely than with the human, so we should kind of roll the dice. We should prefer that dice roll over a human-controlled future.”
You know, there’s a certain logic to it, but I can imagine there’s a part of me intuitively that just is like, “No, let’s just stick with the human that we know,” rather than go with this abstract argument around how maybe a really kind of best-case scenario is more likely with the AI system.
Zershaaneh Qureshi: Same questions to you, Katja.
Katja Grace: Yeah, I agree that we seem to agree a decent amount about what to do about it.
In terms of the weakest parts of my argument, maybe I’m feeling too good about rolling the dice with a human. I think I might not be enough taking into account how likely they are to ask for some sort of narrow thing, or be really foolish quickly and cut off their options in the long term.
And I feel like there was a thing near the start that we maybe disagreed on and didn’t get into that much of like, is there sort of a big space of gradual cases, where some humans take power gradually? I think cases where lots and lots of humans do seem pretty plausible, or maybe that’s like the most likely case of what a non-catastrophic outcome looks like.
What we should actually do [01:17:21]
Zershaaneh Qureshi: Yeah, so you’ve mentioned some disagreements here, and I think something that’s been interesting to me is, despite you having a fair few disagreements on the topics of which of these things is most likely and which of these things would be a bad outcome, nonetheless, you’ve had quite a lot of convergence over what to do.
I’m curious about whether either of you have reflections on what the nature of your disagreement really is, and why it’s not showing up too much when it comes to interventions, and what you think the actual stakes of your disagreement are with that in mind. Should I maybe pass to Katja first here?
Katja Grace: Sure. My guess is that a substantial part of our disagreement is about how likely alignment is, and then I think a reason that we’re maybe not running into disagreements that much about what to do is that the different things we’re concerned about here are like different sides of the same basic problem.
I think a different way of describing our disagreement is: I feel like there’s like a pretty deep general problem where if you make a bunch of creatures that can overpower you and outwit you in every way and put them out in the world, you’re going to run into trouble one way or another. I guess it feels more like, if there’s currently a plan to nuke your city, and you could be like, “Is this person going to press the nuke button, or this person?” Or, “How’s that going to affect the water or the air?” You prefer to just be trying to stop the plan to nuke your city.
I think if you’re worried about powerful AI doing one bad thing or another, it still seems like maybe just good to not build a powerful AI. And my sense is that maybe I feel more like that, because I more think it’s going to be misaligned and that there are just going to be problem after problem. And maybe that Tom feels somewhat more like maybe it will be aligned, and then we might be left with only the problems that are like humans using it for bad — where a human power grab is quite a big one, and maybe you can solve that substantially by focusing on that problem. But yeah, to the extent that that’s also helped by not building powerful AI or doing it more carefully and slowly, that causes us to agree.
Tom Davidson: So I think our disagreements are very much from the super conceptual, abstract side of things. There’s an analogy you can give to deciding what to do and what to believe, which is a kind of a “web of belief” — where you have certain beliefs in the centre, which in some sense are kind of foundational and really important, and they then propagate outwards towards the edge of the web, which is your beliefs about what’s going to happen tomorrow and also your views about what should we do tomorrow.
And I think our disagreements are quite deep into the web — so deep that actually there’s not been much thought about this stuff. And I think, as our discussion showed, there are loads of complicated considerations on both sides. It is quite hard to see how they net out.
But I think you can change the centre of the web a fair bit and still find that when you get to the edge of the web, the action recommendations aren’t that different. It’s like, “This is a super powerful technology which will be capable of disempowering everyone and staging a coup. Man, we need to make sure everyone understands that we need to slow down until we have a plan for how to handle that.” So yeah, not too surprising that a lot of those much more pragmatic opinions are shared.
I do think that if we had really dug into the nitty-gritty of how to implement specific proposals, we might have found that those conceptual disagreements were driving more disagreement.
Like, I would still like more people thinking about superintelligence to focus on the extreme power concentration angle, because it’s like 20 or 30 times fewer I think than thinking about alignment. And that’s only very recently increased, so the historical work is a much bigger ratio.
And in thinking about a centralised project, I would be really upset if we implemented a centralised project and people proposed it without doing the red-teaming for power concentration consequences — where maybe Katja would agree with me that we should do the red-teaming, but it wouldn’t be something that she really cared about.
Katja Grace: Yeah, I also meant to say that what kind of thinking one should do on the margin seems like maybe a clearer disagreement between us. In particular, it seems like probably both of us could work on the things that the other person is working on. And so how much to theoretically think about these things is a disagreement.
What evidence would change their minds? [01:22:36]
Zershaaneh Qureshi: Yeah. I’m interested in what evidence or new information you could imagine getting in the next couple of years that would update you in a different direction from your current position. Are there particular types of warning shots that might make you think differently about this, or particular evidence about the character and behaviours and propensities of AI systems that might skew you one way or another? What are you looking out for?
Katja Grace: I think in practice, if I were to change my mind on this, the evidence seems much more likely to come from me learning more facts about the current world or thinking through it more myself. For instance, I’m assuming various things about what does it look like inside an AI company, and how easy would it be for a particular person to affect how the AI is trained in a sufficiently important way to cause this kind of thing. And if I learned that it more looked like one person could do that sort of thing without many other people having to let it pass, that would make a big difference.
Zershaaneh Qureshi: Same question to you, Tom.
Tom Davidson: I think the big thing is the chance of misalignment. I have updated based on the Hugging Face stuff. I was really surprised by that. And I’m going to keep thinking about it, and I think it’s quite plausible that that shifts me to thinking the misalignment risk is notably bigger overall.
And I think it’ll be really interesting to see the response that the world takes to the recent incidents and any other future incidents that come. I still have a perhaps naively hopeful expectation that, “Wow, we’ve got really quite compelling evidence of messed-up misalignment. Surely as this filters through, it’s not going to happen overnight, but as people can go in with briefings of really compelling evidence, the world is going to change its stance significantly.”
So if that doesn’t happen, which it might not, then again, that’s going to make me more worried about misalignment, because I lose this idea of the warning shots helping out significantly.
And on the power concentration side, there have unfortunately also been updates towards power concentration being more worrying than I thought: actions of the current admin, as well as just it seeming more plausible that there’s going to be a very small number of AI companies, has made those risks seem bigger too. But overall I have been updating towards misalignment recently.
Zershaaneh Qureshi: Yeah, that’s all super interesting to hear. Thank you both so much. I think that’s all we’ve got time for, but it’s been a pleasure having you both on the show.
Katja Grace: Pleasure talking to both of you.
Tom Davidson: Yeah, it was a lot of fun. Thanks very much.
Zershaaneh’s outro [01:25:49]
Zershaaneh Qureshi: All right, that was our first ever debate on the show. I really enjoyed hearing Tom and Katja thinking through these quite thorny questions. And overall I think my main takeaway from this conversation was pretty encouraging.
So it turns out that even if you disagree quite a lot, as Tom and Katja do, over which dangers from AI are most likely and most consequential, then you still might actually converge a fair amount on what we should actually be doing about AI as a society.
And that seems promising, because I guess a worry that I’ve previously had is that there could be pretty big tradeoffs between trying to mitigate the risks of AI takeover versus trying to mitigate the risks of human takeover. And I’d say that I feel a bit less worried about this after hearing Tom and Katja hash out what to do about AI, and having this very honest, reflective conversation, and just hearing them end up agreeing on so much, despite their conceptual differences.
For example, there’s probably a version of a pause in AI development that they’d both at least agree to. And they’re both pretty nervous, maybe in somewhat different ways, about having a single big AI project. And despite his belief that China dominating in AI would be really quite bad from the perspective of power concentration, Tom is still sceptical of arguments that we need to be racing against China. It seemed that both he and Katja maybe favoured negotiating with China about slowing things down versus racing.
At the same time, it does seem like there are still real questions over how we should be allocating very limited resources and things like how much effort should be going into red-teaming a project for the risk of enabling a power grab or something like that.
Anyway, that’s it from me. I really hope you enjoyed today’s episode. Please let us know if you’d like to see more content like this. But you know, if you also never want to see something like this again, I guess that’s also valid. Either way, thanks for watching.