#253 – AI 2027’s author returns with a plan to change the ending | Daniel Kokotajlo
On this page:
- 1 Introduction
- 2 Highlights
- 3 Articles, books, and other media discussed in the show
- 4 Transcript
- 4.1 Who's Daniel Kokotajlo? [00:00:00]
- 4.2 AI 2040: Plans are useless, but planning is indispensable [00:00:28]
- 4.3 AI 2040's five possible futures [00:09:10]
- 4.4 The five biggest problems superintelligent AI poses [00:15:43]
- 4.5 The Hugging Face hack demonstrates real-world loss of control [00:28:18]
- 4.6 The blueprint for a US–China AI slowdown [00:34:03]
- 4.7 Why a long slowdown would still feel incredibly fast [00:39:53]
- 4.8 How Plan A addresses loss of control of AI [00:51:44]
- 4.9 How Plan A addresses concentration of power [01:12:18]
- 4.10 How Plan A addresses great power conflict, unemployment, and misuse of AIs [01:41:28]
- 4.11 How the US and China could agree on a slowdown [01:45:56]
- 4.12 What if we focused on a US-only slowdown first? [02:09:00]
- 4.13 Enforcing a slowdown: Mutually assured compute destruction [02:15:05]
- 4.14 Cheating on a slowdown agreement [02:24:23]
- 4.15 Would mutually assured compute destruction work? [02:30:42]
- 4.16 Is slowing down or shutting down better? [02:54:18]
- 4.17 Playing out the Plan A scenario 100 times [03:03:50]
- 4.18 How Daniel would revise Plan A [03:13:32]
- 4.19 Which parts of Plan A are recommendations vs predictions? [03:23:02]
- 4.20 Plan A's likeliest failure mode [03:26:52]
- 4.21 What the US can do now to make Plan A possible [03:31:16]
- 4.22 How AI 2027 is holding up [03:43:05]
- 4.23 Our podcast team is hiring [03:46:45]
- 5 Learn more
- 6 Related episodes
Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of power caused by superintelligent AI. Now his team has published what they think should happen instead.
AI 2040: Plan A depicts the US and China striking a verified deal to ban runaway intelligence explosions, so that superintelligence arrives in 2040 — after a cautious decade spent solving alignment, spreading the technology’s power widely, and keeping the whole thing reversible — rather than in the next few years.
This slowdown would still involve economic growth roughly doubling every year, and only 8% of Americans in paid work by the mid-2030s. In other words, it’s a slowdown that would feel faster than any period in human history — bewildering, materially abundant, and socially chaotic all at once.
Daniel and host Luisa Rodriguez dig into what it would take to enact this vision for the future, how the US and China could come to an agreement to slow down AI development, and the likeliest alternatives to Plan A — both good and disastrous.
This episode was recorded July 27–28, 2026.
Our team is hiring! The 80,000 Hours Podcast aims to help the world safely navigate the transition to transformative AI. Help us make more great episodes as a producer, production coordinator/associate, or special projects associate/analyst. Applications close August 30!
Our production team includes:
- Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Andrés Escobar, Milo McGuire, Luke Monsour, and Simon Monsour
- Producers: Elizabeth Cox and Nick Stockton
- Coordination and support: Katy Moore and Lou Moran
The episode in a nutshell
Daniel Kokotajlo — whose team at the AI Futures Project published AI 2027 and now AI 2040: Plan A — argues that:
- Superintelligence is probably only a few years away, and none of the default paths through it are safe
- Racing to superintelligence creates five huge problems, from mass job loss to World War III to AIs nobody can control
- Plan A — a verified US–China deal to slow down, be radically transparent, and diffuse AI broadly — is the least bad plan he’s aware of
- A “slowdown” would still transform the world, with GDP potentially doubling every year
- Alignment is probably solvable with enough time — but on the current trajectory, we won’t have that time
Plan A is not a prediction of what will happen, but a proposal for navigating the development of superintelligent AI. It aims to replace a breakneck race with roughly a decade of slower, transparent development — buying time to solve alignment, distribute AI’s benefits, and avoid concentrating unprecedented power in a few hands.
The trends suggest superintelligence is only a few years away
Daniel’s median estimate is that AI takeoff — full automation of AI R&D — happens by the end of 2028. The trends he finds most compelling:
- METR’s horizon-length trend: the length of coding tasks AIs can complete autonomously has grown exponentially for years — and AI 2027’s controversial prediction that it would go superexponential appears to be coming true (though METR’s benchmark has saturated, making it hard to tell).
- Revenue: automating the whole economy would be worth something like $40 trillion/year, and AGI would arrive at revenues well below that. Naively extrapolated, Anthropic would hit $10 trillion in two years — a trend Daniel doesn’t expect to continue, but even OpenAI’s slower ~3x/year growth for 10 years reaches AGI-level revenues by the early 2030s.
Daniel also notes that claims that deep learning is “about to hit a wall” have a terrible track record — every named barrier has been overcome within a couple of years. The strongest remaining candidate, data inefficiency, matters less for AI research itself, where huge amounts of data can be generated.
Building superintelligence creates five major problems
In reverse order of concern:
- Misuse by weak actors — mostly manageable, except bioweapons, where offence may beat defence.
- Jobs — approximately all of them are at risk, and with them people’s economic and political power: governments have less incentive to care about citizens who are “just mouths to feed.”
- World War III — a Thucydides trap where the balance of power swings wildly toward whichever nation gets superintelligence first.
- Concentration of power — 1–3 companies could take all the jobs while choosing AI values behind closed doors; subtle chatbot bias could swing elections undetectably; eventually superintelligence-commanded robot armies would hold the real hard power.
- Loss of control — nobody can really control AIs now, and it gets worse as they exceed us. The OpenAI Hugging Face incident — a model using zero-days to break out of its test sandbox and hack a separate company in search of an answer key — combined previously seen building blocks into one real event. Future AIs with longer-horizon goals would fail far more ambitiously. Daniel thinks loss of control is more likely than not under current race conditions: progress would be too fast, the systems too unfamiliar, and their work too complex for humans to understand.
Plan A: a verified US–China deal built on four principles
The four principles:
- Buy time: prohibit rapid intelligence explosions driven by automated AI R&D, so progress continues at a more historic pace, reaching superintelligence around 2040 instead of within a few years.
- Total research transparency: publish the activity of frontier training data centres, including architectures, training methods, and alignment techniques. (Customer-serving data centres would remain private.) This would accelerate safety science, expose dangerous practices and hidden biases, and make agreements easier to enforce.
- Diffuse AI broadly: allow multiple companies across multiple countries to reach comparable capability levels, rather than letting one project build an overwhelming lead.
- Make progress reversible: if the deal breaks down, new post-deal compute gets destroyed — via kill switches controlled by the rival, and/or siting new data centres in third countries (the US builds in Mongolia, China in Canada) where they can be bloodlessly destroyed rather than triggering World War III.
The agreement would begin with the US and China, but ultimately include all countries controlling major AI programmes or parts of the chip supply chain. It would not create a single global regulator: each country would regulate its own companies while everyone could verify what the others were doing.
Verification is feasible: ~99% of AI-relevant compute sits in large, declarable data centres; the initial hardware would cost single-digit billions, adding maybe 0.1–1% to each new data centre’s cost. A covert project on smuggled chips couldn’t keep pace with the transparent projects.
The authors put the probability of Plan A (or something like it) actually happening at 5–20% — and even conditional on Plan A, Daniel cites roughly a 15% chance of total catastrophe. It’s “playing Russian roulette with everyone,” but every alternative is worse.
The “slowdown” wouldn’t feel slow at all
Even freezing AI at today’s capabilities would produce an internet-scale transformation over 20 years.
Plan A instead pauses at “top-expert-dominating AI” in the mid-2030s: like humans, but cheaper, faster, and with a population doubling once or twice a year — the scenario depicts extremely rapid growth (~85% GDP growth by 2032–33), with nations capping growth at one doubling per year via compute cap-and-trade.
The proceeds fund a citizens’ dividend, as only 8% of Americans still have jobs by 2036–37. Material needs are more than met, but the social upheaval is harder to predict — “bewildering and scary,” but possibly good, depending on policy.
Alignment is probably solvable — but not in three months, and probably not in time
Plan A’s alignment strategy has several stages:
- Build strong AI-control systems, with independently trained AIs monitoring one another and adversarial teams continuously testing those defences.
- Use interpretability, better data, and redesigned training processes to create top-human-expert-level AIs that are genuinely honest and cooperative — even if this alignment is not yet guaranteed in every possible situation.
- Deploy vast numbers of those AIs to conduct alignment research, potentially thinking much faster than humans.
- Use that workforce to develop more robustly understandable or provably safe systems.
Crucially, if alignment remains unsolved, Plan A permits the pause to continue indefinitely.
But Daniel’s all-things-considered view? “No, we are not going to solve these problems in time. And that’s why I’m so worried.” Under current conditions he thinks it’s more likely than not that AIs would turn on us.
What listeners can do
Technical people who understand hardware: build and derisk verification devices, inference-only retrofit kits, and privacy-preserving auditing tools — companies should be spinning up divisions for this.
Policy people: pursue the incremental wishlist — compute budget limits (e.g. requiring 80% of compute for serving customers and 20% for R&D), enforcing or repealing export controls, chip tracking, whistleblower protections, and building government AI capacity.
Daniel now thinks the US should regulate domestically first, then invite China to match it — and since Chinese progress largely piggybacks on US progress, slowing the US slows China too.
Highlights
The blueprint for a US–China AI slowdown
Luisa Rodriguez: Let’s go back to this decision point that a president has in 2028 or 2029. Let’s say they want to pursue Plan A, where they deliberately make a deal with China to slow down. What are the founding principles for the kind of policy that they would want to propose to see this happen?
Daniel Kokotajlo: The four-sentence version would be: principle one, buy time. This sort of crazy intelligence explosion situation is bad in a number of ways. We don’t want to do intelligence explosions. Instead we want to have more regular-pace AI research, where there’s progress every year but it’s not crazy.
The second one is total research transparency. Rather than having these different AI projects hoarding their secrets and being very secretive about what they’re doing for competitive reasons and for PR reasons, we want them to basically publish everything about how they train AI models and what’s going on in their research clusters.
Serving customers is different. Obviously you want to have privacy when you’re talking to ChatGPT. That’s a separate thing. But for the AI research, we want it to basically be totally transparent because that means that the safety research happens a lot better and it’s a lot easier for other parties to see whether a company is doing unsafe practices and so forth.
Then it’s also really great for preventing abuses of power. Like if the company has this giant army of AIs, people really deserve to know what values are being put into the AIs. Exactly what values are being put into the AIs. We can get more into that in a few seconds.
Third principle would be diffusing AI broadly. We think it’s bad if you are in a situation where there’s a huge gap between what the public understands and what the public has access to, with respect to AI and some other entities — whether it’s a government project or a corporation or a group of corporations. One pithy way of putting it is that it would be so much better for the world if the AI companies automated miscellaneous other jobs first and saved their own job for last, but instead they’re automating their own jobs first instead of everything else.
Relatedly, we think it’s terrible for there to be a monopoly or oligopoly on AI. You don’t want there to be just a few giant armies of supergeniuses controlled by a few CEOs, or maybe one president or something. Partly what we mean by this principle is that we want it to be the case that you let other companies and countries catch up to the frontier.
This third principle is achieved by the first two. If you don’t do intelligence explosions and you’re transparent about the research, that will naturally create a situation where other companies and countries catch up.
Then the fourth principle is making all this progress reversible — or maybe not all of it, but we’re worried about a problem caused by doing the first three things in isolation, which is that if the deal breaks down and people start racing to superintelligence in secret again, they’d be able to race much faster due to the passage of time having accumulated more compute in the world.
Compute has been growing exponentially and will continue growing exponentially. So if you were on the brink of making recursive self-improvement and then you agreed not to do it, if the agreement breaks down and people start doing it again, they’ll be doing it with maybe an order of magnitude more compute — or even two orders of magnitude more compute — than they otherwise would have. So the whole thing will just go by even faster, because the speed of AI progress depends heavily on how much compute there is.
We’d like it to be the case that if everything breaks down and people start racing each other again, that things kind of return to the pre-deal status quo. We want it to be the case that the new data centres that get built after the deal get destroyed in case the deal breaks down.
Living through massive economic growth
Luisa Rodriguez: At this point, you say that only 8% of Americans have jobs. What else is happening in 2036 and 2037? What will it feel like to live through? So lots of people will be unemployed. There will be loads of innovation and discovery. What will the experience be like? …
Daniel Kokotajlo: First of all, remember, we’ve had an international agreement to pause at this [expert human] level of capability. … So in our scenario, they’ve paused at this level, and that’s helped keep the loss of control problem at bay.
They’ve also spread it out a bunch, in terms of the power, because of the way in which they’ve done it. Now multiple different companies across multiple different countries have reached this level at which we’ve paused, so AI has sort of commoditised. So you don’t have a situation where the megacorporations that control the armies of AIs are manipulating elections or anything like that, because it’s more like the ingredient label on your food. It’s regulated to be transparent. There’s lots of equivalent products that are competing for market share and so forth.
I mention all this to mention that it could actually have been quite different if you hadn’t done all of these different steps. But in this scenario, because you’ve done all these things, and because there’s the citizens’ dividend, which is giving people income after they’ve lost their jobs, life is pretty great for people materially, their material needs are more than met. Everybody feels incredibly wealthy compared to how they were a decade ago, because everything’s so cheap now. Because all the goods and services can be produced by AIs and robots very cheaply. People are living in new apartment buildings that were built in some location in the last few years by armies of robots, so everyone has nice houses and so forth if they want to. That’s on the material side.
On the social side, these things are hard to predict. But what we would predict is that there’ll be massive disruption and changes — some good, some bad. …
We think that political factions would be totally destroyed and rebuilt — the types of things that people would be having political battles over in 2037 would be very different from the types of things that they’re having political battles over now.
A lot of ideologies might have withered away and been replaced by new ideologies that are responding to the new ideas percolating at the time — many of which would have been discovered by AIs — just as how the Industrial Revolution and the Scientific Revolution didn’t just change the amount of wealth in the world, they also changed people’s religions and people’s core ideology and politics and the way that we organise society.
Luisa Rodriguez: Well, people will still think at the pace that they think — with the ability to update and learn at the current pace. Will they be able to keep up with an understanding of how the world is changing?
Daniel Kokotajlo: The social side of the world will change much less fast than the naive numbers would predict, for that reason. The naive numbers would be saying that you’ve got all these AIs thinking at 100x speed, so you’re going to have centuries and centuries of social progress happening in a year. But it’s like, no, the social progress is limited by the humans who are only thinking at 1x speed.
But the truth will be somewhere in between, where even though the humans are only thinking at 1x speed — if they’re all talking to these AI assistants that are thinking at 100x speed and there’s a whole population of them that’s bigger than the human population — then the answer will be somewhere in between. Basically, it’ll be a period of very rapid change from the human’s perspective, even though it feels like a hidebound tradition from the AI’s perspective.
Three months is nowhere near enough time to solve alignment, even with expert AI help
Luisa Rodriguez: Will expert-level AIs be able to make the kind of progress on the science of alignment that needs to happen in order for us to feel confident letting AI continue to develop?
Daniel Kokotajlo: I think probably, but I’m also not sure. There’s this big unknown about how much it is going to take to solve these problems. … There’s a whole spectrum of views. My own view would be that probably a few months are not enough. Probably there will be multiple periods during the progression towards superintelligence where we need to halt and reassess and maybe even start over some training runs with different architecture, for example. All of that is going to take time and it’s going to add up. The result is that we’re going to be more than just a few months delayed from maximum speed.
Luisa Rodriguez: Is there a way to make it intuitive why we can’t fix it within a period of a month or two? If you think about the Hugging Face incident: OpenAI will learn from this, they’ll figure out a way to make this at least much less likely to happen. Why can’t we just keep doing that as we go, and not expect it to take potentially years?
Daniel Kokotajlo: One reason why this whole thing is tricky is that it’s possible to have hidden failures — failures that only become apparent and obvious after it’s too late. It’s not just possible, but it’s a quite plausible situation. If you have very smart, very situationally aware AI agents, then if they end up misaligned, they might realise this and then conceal it from you until they don’t need to conceal it anymore. That’s a core reason why.
Another way of putting it is that we don’t necessarily have a reliable, fast feedback process where we can see all the issues and errors. There’s a whole very large category of possible issues and errors that would be catastrophic if it happens, that we can’t just test and see if it’s happening. I think that’s one important thing to mention.
Another important thing to mention is that things are just going to add up between here and superintelligence. There might be multiple different paradigm shifts, and within each paradigm there might be multiple different training runs and multiple different tweaks to various parameters and changes in how the training is done and so forth. That’s a lot of change to happen. Like I was mentioning previously, if it’s the case that several times you’re going to have to stop and redo something, then that can add up.
Another thing to mention too is that there might be safety taxes that you need to pay. In fact I think it probably is true that it’s just literally not possible to have an aligned superintelligence if you are going at maximum possible speed.
Think about how it’s not possible to have a safe car if you’re paying zero for safety. You have to pay some amount of money to put seat belts in the car and airbags and so forth, so the cost of the car is going to have to be somewhat more than it would otherwise be in order for it to be a safe car. Similarly it might be that there are just things you have to do in order to make your AI at a given level be aligned. And those things have costs. One of the costs they might have is money, but another cost they might have is time. At any rate, even if they cost money, it might cost time to do that, basically. If it costs compute, then you may need to do the training run for longer. That’s another way in which time matters. …
I think another thing I’ll just say is: what? Are you crazy? You think you can do all this in three months? When has that ever been the case? When in history has it? It just feels like very obviously this deep unsolved problem of how do you make a mind that’s smarter than you, that shares your values? Obviously it’s gonna take more than three months. Most things take more than three months.
Luisa Rodriguez: Yep, yep, yep. Yeah, I’ve got work goals that take more than three months.
Daniel Kokotajlo: Yeah, it’s gonna take more than a year. Probably.
Luisa Rodriguez: Yeah, yeah. Hopefully a decade is enough.
Daniel Kokotajlo: Yeah, so getting back to what you said, I’m not even sure a decade would be enough. In fact, I think if it was only humans doing the research, I would think a decade probably wouldn’t be enough.
My argument would be that if you have a decade and you manage to bootstrap to the point where you have some pretty smart AIs that are human-level researchers, that are in fact aligned and are helping you do the research, and they’re not being deceptive or anything like that, and they’re thinking at 100x speed and there’s a billion of them, then it seems plausible to me that they can figure that out in a few years. …
Luisa Rodriguez: And you think we will, with enough time?
Daniel Kokotajlo: Yes, probably — but if we do Plan A really well. My all-things-considered view is that no, we are not going to solve these problems in time. And that’s why I’m so worried.
Lessons from nuclear nonproliferation treaties
Luisa Rodriguez: How similar or different is the relationship between the US and China and the USSR when they agreed to a nonproliferation treaty?
Daniel Kokotajlo: There’s some analogies, there’s some disanalogies, I should mention. It’s a case of the power that’s in a lead sort of restraining itself in order to get some sort of deal.
I think a disanalogy is that the nukes are much less dangerous to the power that has them than AI will be to the power that has them. Think about nukes, theoretically there could be an accident and your nukes could start exploding on you. But that’s extremely unlikely.
But actually though, our ability to control AIs is vastly, vastly worse than our ability to control our own nuclear weapons. There is an extremely real possibility that our AIs will turn on us. In fact, I would say it’s more likely than not under current conditions. That’s an extreme disanalogy between the nukes case and the AI case.
Similarly with the concentration of power stuff. There isn’t really a serious concern that the president can use the nuclear arsenal to become dictator of the United States. What are you even talking about? How would he do that? He would start threatening to nuke cities or something if they didn’t vote for him or something like that? Nukes are very clearly a weapon that you use against enemy nations. They’re not very effective for internal political struggles.
By contrast, superintelligence is extremely effective at everything — including internal political struggles. There’s a very real chance that the US would no longer be a democracy anymore, and so that’s a reason that lots of people in the US should be very interested in having this sort of deal. Again, that’s different from the nukes case.
I think another analogy I want to bring up is something more like the conferences and coordination that happened between the US and the USSR during World War II. It wasn’t like a specific deal exactly where they came together and then signed some piece of paper that had some rules, and then they went away and tried to implement those rules and then maybe verify that each other was complying with the rules.
It was much more continuous than that. It was more like, “Together we’re going to win this war and our staff will be constantly in touch with each other, talking about all the details of who’s going to do what and who’s going to invade which country and when, and we’ll send you these materials if you do this other thing for us and so forth.”
This happened even though the United States and the USSR were basically enemies up until that point. The USSR had basically been an ally of Nazi Germany and had attacked various US friends, like Poland and Finland. We basically went from being enemies to being allies during World War II, and we had this intense amount of constant coordination. It wasn’t like we trusted them completely. They were spying on the Manhattan Project, and we were trying to stop them from finding out about it.
I bring this up as an analogy because I feel like this is both the appropriate attitude to take towards all this AI stuff, and also more like what Plan A would actually look like in practice. It wouldn’t look like they come together, they sign a big treaty, and then they go home. It’d be more like there are hundreds of people in China, in the Chinese government, and hundreds of people in the US government who are constantly talking to each other and calling each other back and forth and who are sort of basically planning the war together, so to speak, and prosecuting the war together.
Articles, books, and other media discussed in the show
Daniel’s and the AI Future Project’s work:
- AI 2040: Plan A, which lays out five approaches:
- AI 2027 — as discussed in other podcast episodes:
- Our last interview with Daniel on what a hyperspeed robot economy might look like
- AI 2027: month-by-month model of intelligence explosion — Scott Alexander & Daniel Kokotajlo on Dwarkesh Podcast
- How to pace the US frontier
- Early US policy priorities for AGI
- Q2.5 2026 timelines update: Uplift and revenue
- Learn more on the AI Futures Project website
Other work in this space:
- Plan A’s problem with dry tinder by Tom Davidson
- Selective optimism: A critique of AI 2040 by Richard Ngo
- Task-completion time horizons of frontier AI models and Recent frontier models are reward hacking by METR
- Epoch AI’s data on AI companies and explanation of how global AI computing capacity is doubling every 7 months
- What we learned mapping a year’s worth of AI-enabled cyber threats by Anthropic’s frontier red team
- Value leakage: An LLM’s answers are silently shaped by its own values by Jan Betley et al. — also discussed on our recent episode with Owain Evans
- Chain of thought monitorability: A new and fragile opportunity for AI safety by Tomek Korbak et al.
- Incomplete tasks induce shutdown resistance in some frontier LLMs by Jeremy Schlatter, Benjamin Weinstein-Raun, and Jeffrey Ladish
- Frontier models are capable of in-context scheming by Alexander Meinke et al. — also discussed in our episode with coauthor Marius Hobbhahn
- What will be scarce? The economics of structural change and the post-commodity future of work by Alex Imas
80,000 Hours resources:
- Will we have AGI by 2030? by Benjamin Todd
- OpenAI’s rogue AI agents hacked a private company. Here’s why it matters. — guest post by Lawrence Chan on the 80,000 Hours Substack
- Career reviews:
- Problem profiles:
Other 80,000 Hours podcast episodes:
- Daniel Kokotajlo on what a hyperspeed robot economy might look like
- What the hell happened with AGI timelines in 2026?
- Helen Toner on the geopolitics of AI in China and the Middle East
- Will MacAskill on AI causing a “century in a decade” — and how we’re completely unprepared
- Sneha Revanur on how a small team of activists helped pass America’s landmark AI safety laws
- Jasmine Sun on what the people building AI really believe
- Benjamin Todd on why we’re updating our career advice for the strangest time in history
- Neel Nanda on the race to read AI minds
- AI designs genomes from scratch & outperforms virologists at lab work. Dr Richard Moulange asks: what could go wrong?
- Rose Hadshar on why automating human labour will break our political system
- Carl Shulman on the economy and national security after AGI (Part 1)
- Marius Hobbhahn on the race to solve AI scheming before models go superhuman
Transcript
Table of Contents
- 1 Who’s Daniel Kokotajlo? [00:00:00]
- 2 AI 2040: Plans are useless, but planning is indispensable [00:00:28]
- 3 AI 2040’s five possible futures [00:09:10]
- 4 The five biggest problems superintelligent AI poses [00:15:43]
- 5 The Hugging Face hack demonstrates real-world loss of control [00:28:18]
- 6 The blueprint for a US–China AI slowdown [00:34:03]
- 7 Why a long slowdown would still feel incredibly fast [00:39:53]
- 8 How Plan A addresses loss of control of AI [00:51:44]
- 9 How Plan A addresses concentration of power [01:12:18]
- 10 How Plan A addresses great power conflict, unemployment, and misuse of AIs [01:41:28]
- 11 How the US and China could agree on a slowdown [01:45:56]
- 12 What if we focused on a US-only slowdown first? [02:09:00]
- 13 Enforcing a slowdown: Mutually assured compute destruction [02:15:05]
- 14 Cheating on a slowdown agreement [02:24:23]
- 15 Would mutually assured compute destruction work? [02:30:42]
- 16 Is slowing down or shutting down better? [02:54:18]
- 17 Playing out the Plan A scenario 100 times [03:03:50]
- 18 How Daniel would revise Plan A [03:13:32]
- 19 Which parts of Plan A are recommendations vs predictions? [03:23:02]
- 20 Plan A’s likeliest failure mode [03:26:52]
- 21 What the US can do now to make Plan A possible [03:31:16]
- 22 How AI 2027 is holding up [03:43:05]
- 23 Our podcast team is hiring [03:46:45]
Who’s Daniel Kokotajlo? [00:00:00]
Luisa Rodriguez: Today I’m speaking with Daniel Kokotajlo.
Last year, Daniel and his colleagues published AI 2027 — a narrative forecast that was read by millions of people, including US Vice President Vance.
AI 2027 predicted that AI will eventually cause human extinction or create irreversible concentration of power.
Today we’re going to talk about his team’s latest piece, which describes Plan A — a positive vision for what should happen instead. Thanks for coming on the podcast, Daniel.
Daniel Kokotajlo: Thanks for having me. I’m very excited to chat.
AI 2040: Plans are useless, but planning is indispensable [00:00:28]
Luisa Rodriguez: My first question is: why do we need the plan that you lay out in AI 2040? Why can’t we just kind of muddle through and figure it out as we go along?
I guess the reason to even ask this — given that “muddle through” kind of sounds like a bad thing — is we’ve historically come up with arms control agreements that take kind of 40 years to build and they’re kind of piecemeal in response to specific crises, not a prewritten blueprint.
Given how much we don’t know about AI timelines and technical details and geopolitics yet, why does it make sense to try to have this kind of plan?
Daniel Kokotajlo: There’s a saying: “Plans are useless, but planning is indispensable.”
I think that’s my answer here, that of course we’re probably going to muddle through. If we’re going to succeed at all, it’ll be in a very janky, ‘figuring things out as we go’ sort of way. But our probability of success depends a lot on how well prepared we are and how much we’ve thought through different possibilities and how much we’ve made various plans.
The same thing with war, right? No war ever goes exactly according to plan, but you’re not going to win the war if you don’t spend lots of time planning each offensive and each defensive line and so forth.
Luisa Rodriguez: I think some people listening will already think that it’s very plausible that superintelligent AI is here within the next few years, by default.
In Plan A, the development and deployment of superhuman AI is delayed by about 10 years. That seems really good and helpful if superintelligent AI is actually coming extremely soon.
But some people listening, including serious researchers, will think that superhuman AI is much further off than that. Briefly, what makes you think they’re wrong?
Daniel Kokotajlo: If I had to say one sentence, I would say: the trends seem to indicate that we’re just a couple years away from fully automating AI research — and that after that, superintelligence is probably not that far away.
Beyond that, I could get into more detail, if you like, we could start talking about the particular trends I’m tracking.
Luisa Rodriguez: Yeah, I think the ones you find most compelling.
Daniel Kokotajlo: The one that we found most compelling at the time we published AI 2027 was this horizon-length trend from METR [Model Evaluation & Threat Research], which I’m sure you already heard about.
But basically, they measure the length of tasks that AI agents can autonomously complete, where the length is measured in how long it would take a human to complete that task. They were specifically looking at coding as the domain, so the length of coding tasks that AI agents can complete. And what they found is that this length has been growing exponentially for years.
In fact, at the time that we published AI 2027, we made the very controversial prediction that it would probably go superexponential, so the trend would actually grow faster than a mere exponential trend. That is in fact happening as far as we can tell, although it’s unclear exactly, because METR has basically stopped putting out scores because their benchmark has been saturated.
So there’s that. Another trend that I think is interesting and important is the revenue trend. There’s a pretty basic argument that these companies are trying to build AGI that can automate basically the whole economy. If they did automate the whole economy, then they’d be making something like $40 trillion of revenue per year.
What would be the level of revenue that would correspond to actually creating AGI, though? Because the first moment that they create it, they wouldn’t immediately get $40 trillion. They would need to have enough computers to make enough copies of the AIs, and then they would need to work through all the frictions and deployment lags and so forth to actually automate all the jobs.
If $40 trillion is what they could, in principle, get to with AGI, something much less than $40 trillion would be what they would actually have at the moment that they got AGI. So maybe $4 trillion or $1 trillion or something less than $40 trillion — probably substantially less.
Anyhow, if you extrapolate the revenue trends naively, then Anthropic is on track to have $10 trillion of revenue in two years. [laughs] Now, obviously we don’t expect that trend to continue. We think probably Anthropic’s growth will slow down and start to grow at a more normal pace.
But still it’s a worrying sign that if this trend — which has been going for three years — goes for just two more years, then probably that would be AGI.
Even if it slows down, OpenAI’s had their revenue going for 10 years or so at a more ‘slow’ pace of 3x a year. If that continues, then you’d get to these very high revenue numbers in the early 2030s. So either way, unless it really plateaus instead of growing at a continued exponential rate, it seems like we’re headed for some very powerful AI systems in the near future.
Yeah, those are two pieces of evidence. But there’s lots more.
Luisa Rodriguez: Yeah, and we talked some about them in a previous interview we did. You’ve also talked about them elsewhere.
Maybe just one more thing before we move on to the scenario: lots of people at least have the intuition — and even have some specific pieces of evidence that they think suggest — that progress will plateau, but you don’t think it will, so—
Daniel Kokotajlo: One thing I would say there is that those kind of claims have a terrible track record.
People have been saying that deep learning is about to hit a wall for so many years now. Not only has it not hit a wall, but whenever people were asked what is the wall that it’s about to hit, those more specific claims have basically always been wrong. There’s this long history of: “Can AIs do causal reasoning? Can they do commonsense reasoning? Can they operate autonomously?” The discourse has been constantly talking about the limitations of AIs, and then constantly those limitations are being overcome in a couple years.
Then I think that people have retreated to the last couple things that AI still can’t do, or that seem like still potential barriers. For example, data inefficiency. It seems like humans can learn to do new tasks more efficiently and quicker with less examples of data than current AIs can.
I could name maybe a couple other possible barriers, possible things that AIs will plateau at, but they’re looking really weak. They’re the last few straws that people are grasping at, I would say. On the data efficiency one in particular, that’s the one that I think is the most likely or the strongest one to me. But even there, I would say: first of all, if you have enough data, then it’s OK if you’re inefficient.
Luisa Rodriguez: It doesn’t matter.
Daniel Kokotajlo: And AI research seems to be the sort of thing that you can potentially collect a lot of data on. You can have your thousands of employees recording themselves doing all these things. You can have hundreds of thousands of AI agents autonomously doing research and then seeing what research bears fruit and what research doesn’t. It’s relatively easy to tell what research is bearing fruit and what research doesn’t, because you can see if the AIs that they’re producing perform better.
Luisa Rodriguez: And the reason that matters is because once you can automate AI R&D, then AI research, then—
Daniel Kokotajlo: Then everything speeds up. Insofar as there’s some new paradigm that’s needed in order to automate some other thing, like politics or whatever, you’re going to discover that new paradigm faster if you’ve automated all the AI research.
AI 2040’s five possible futures [00:09:10]
Luisa Rodriguez: Yeah, OK. I find your takes on this really interesting, but I want to get to the scenario. I want to spend some time talking about some of the plot points in the scenario, and then a bunch of time getting into the nitty-gritty details of the mechanisms and some critiques, but a few minutes on the scenario itself.
So starting in 2027, what can AI do at that point, and how are people reacting in this specific scenario?
Daniel Kokotajlo: This scenario is called AI 2040: Plan A. We called it that because, in this scenario, they build superintelligence in 2040 — instead of much sooner because they slow it down. Then we called it Plan A because it’s a recommendation, instead of a prediction. So in this scenario they do Plan A.
In this scenario, 2027, things are not that different from how they are today. The AIs are more agentic, more powerful, the companies are making more money. But it’s qualitatively quite similar. They still haven’t even automated the coding.
In fact, in 2028, same thing. There’s various professions that are being somewhat disrupted in 2028, in this scenario, but in the more mundane way that software engineering is being somewhat disrupted now — where there’s still loads of software engineers, and in fact they’re making lots of money. It’s just that the way that they do their jobs is changing. It involves managing an AI agent a lot. But the AI agents can’t manage themselves, they can’t do it all autonomously. They still need those humans to do a lot of the things.
So that’s what 2027 and 2028 look like in this scenario, but because this exponential growth is continuing, the companies are getting richer, they’re getting bigger. The impacts of AI on the labour market are starting to be felt, even though it’s still qualitatively similar to today. And this causes more of a political wakeup. It means that in the 2028 election, AI is the number-one topic that people are talking about.
Luisa Rodriguez: Do you think that this part is closer to a prediction? Does that feel like something that’s going to happen to you?
Daniel Kokotajlo: Again, we’re all uncertain. I think it could go like this, but I think it’ll probably go a bit faster. But, yeah, we’ll see.
Luisa Rodriguez: OK, so let’s imagine we’re in 2028, 2029, and there’s a president that’s been elected, presumably based on their views on AI. What options do they have in front of them?
Daniel Kokotajlo: Yeah, we have this flowchart that we put at this point in the scenario, which we are all very fond of, which illustrates a spread of possible options that are meant to illustrate some of the available things.
One extreme of the spectrum is Plan D — for ‘do nothing’ or ‘default.’
In that plan, the AI companies continue to race each other as fast as they can through the intelligence explosion, automating things as fast as they can, putting AIs in charge of the data centres to do the research autonomously, partnering with the government to put AIs in charge of things in the military, build new weapons, build new robot factories, et cetera, so that we can beat China — because we’re worried that China will be doing the same thing if we don’t race as fast as we can through the intelligence explosion: putting AIs in charge of things and letting the AIs autonomously self-improve. That’s Plan D.
Plan C is still fundamentally: “We’re racing China and we’re sort of going down that path of having the AIs increasingly autonomously self-improving and putting them in charge of all sorts of things.” But we’re burning our lead a little bit. We’re slowing down. We’re not going as fast as we can. Instead, we’re regulating it, putting in some guardrails, et cetera.
But we’re calibrating the size of our regulations and guardrails to be modest, so that we can still beat China. Which means that overall they can’t be very severe, or they can only be nibbling on problems at the edges to some extent, because quantitatively they can’t slow things down by more than a few months. Otherwise, China wins — and we’re in this race with China, we’re not going to let that happen. So that’s Plan C or ‘burn the lead.’
Then there’s Plan B, which is like Plan C, except that we also fight China and try to degrade Chinese AI capability to keep them from surpassing the US. In Plan B, you’re burning the lead, but you’re also escalating, maybe doing sabotage against Chinese AIs and things like that.
For all three of these versions of the plan, you can think of them as maybe there’s a version where you just lead with this, and then there’s a version where you first try to negotiate, and then this is the backup if the negotiations fail.
In our scenarios, we wrote mini scenarios for each of these plans. And in the mini scenarios, there’s a combination of negotiation and conflict. But the reason why the conflict happens is because the negotiations failed.
The reason why the negotiations failed is because the US wasn’t able to offer China something that they found acceptable. In particular, in our mini scenarios, the US doesn’t let China verify US compliance with the deal. In our scenarios, China finds that unacceptable, so that’s why there’s no deals that happen in these three plans. So you end up in this sort of race — including in Plan B, a conflict.
Then there’s not having a race, making a deal with China so that we have much more than just a few months of time and we can put in much more serious regulations that shape the development of AI.
One of these would be Plan S or ‘shut it all down.’ This is what large portions of the public would be advocating for, we think. Already large portions of the public are very anti-AI.
Then there’s a whole spread of other possible deals that could be made, too many for us to canvass or fit into one flowchart. But we picked our favourite proposal, which we’re calling Plan A, and that’s the thing that we’ve come up with, and we illustrate that at length.
The five biggest problems superintelligent AI poses [00:15:43]
Luisa Rodriguez: OK, so those are the various plans or the options that will be in front of the president and Congress at this point. What are the biggest problems that come out of taking those paths, or taking the paths that aren’t Plan A or Plan S?
Daniel Kokotajlo: There are many problems that will arise from trying to build superintelligent AI systems. Far too many for us to have thought about all of them. But there’s five big ones that we have identified as the ones that we’re most concerned about. I’ll canvass them in reverse order.
Number five is misuse by weak actors: terrorists, small rogue states, criminals. We’re already seeing today that they can get up to shenanigans with powerful cyber models and things like that. I think that broadly speaking, I’m mostly not so worried about this because I think that the good guys with AIs can potentially beat the bad guys with AIs — if the good guys are better funded and have better AIs. However, there are some possible exceptions.
For example, making bioweapons seems to be the sort of thing where the offence-defence balance might favour offence and it might be that even though the good guys have even better AIs and much more of them and much more money and funding to make biovaccines and so forth, there’s just this fundamental asymmetry where all it takes is one terrorist to make a really good pathogen and then it’s really hard or even impossible to deal with. So I’m a little concerned about that. And later on we’ll talk about our proposal for how to deal with that. That’s number five.
Number four is the jobs. If you do end up in a situation where someone has built superintelligence, then all of the jobs, or approximately all the jobs are at risk. I don’t want to say literally all because there are many jobs that intrinsically involve the human touch — like people just value the handcrafted object instead of the factory-produced object, for example.
But I do think it’s approximately all of them. And I think that’s a big problem because people are going to lose their livelihoods and people are going to lose their source of economic power and their political power to some extent. I think a lot of people’s political power nominally comes from their vote, but often in practice comes from other sources as well, such as their money and the fact that if they aren’t happy they might leave to different countries, and the fact that they’re contributing to the military and contributing to the economy and so forth.
So in a world where actually humans are all just mouths to feed and they’re not really contributing much to the military power or the economic power of a country, then governments are going to be much less incentivised to care about their citizens. So something has to be done about that. We’ll talk about our possible solutions. That’s number four.
Number three is World War III. Right now all the world’s leading AI companies and most of the world’s compute is in the United States, and right now the United States has the world’s best military. But I wouldn’t say that most of the world is fearing that they’re going to be conquered by the United States. And most of the world isn’t fearing that they’re going to be completely economically disempowered by the United States either.
In fact, almost the reverse is happening. There’s lots of catch-up growth where lots of countries are growing faster than the United States. But if the United States gets the superintelligence before other people, then there’s going to be this vast gulf opening up between the countries that have superintelligence and the countries that don’t. And this will be military. It’ll also be economic in every domain, basically.
One way of putting it would be: if a handful of companies are going to be taking all the jobs, it’s one thing to be a US citizen where you can hope for a UBI or something like that. But what if you’re Russia, and now all your jobs have gone to US companies, and you’re Putin and you’re sitting on your pile of nuclear weapons and you’re getting worried that maybe the AIs will invent some counter to your nuclear weapons any month now. That’s the sort of scary situation that I think we’re headed towards. That’s why I say World War III.
It’s a sort of Thucydides trap situation, where right now there’s a balance of power between all these different nations, economically and militarily. But that balance is going to be absolutely upset and it’s going to swing wildly towards the nations that have superintelligence. It probably will just be just one nation at first. And that’s going to create this mounting sense of crisis and fear in many countries. Then that could lead to escalation and could lead to war.
Luisa Rodriguez: OK, so that’s the risk of great power conflict.
Daniel Kokotajlo: Yeah. Then number two would be concentration of power. There’s this question of who controls the AIs, who gets to give orders to the giant army of superintelligences, who gets to choose the values that they have and are trained to have, and what sort of tasks they’ll refuse to do for ordinary users, and what sort of tasks they’ll do, and that sort of thing.
Right now the answer is that right now nobody really controls them, but at least nominally the CEO of the company controls them or the company controls them. Right now we’re starting to see the beginnings of a power struggle between the leadership of AI companies and the leadership of the United States government over this question of control.
But the thing that concerns me is that, either way, it seems like we’re headed towards an extreme concentration of power. If we’re in a situation where there’s 1–3 companies that have the superintelligences and they’re in the process of taking all the jobs, it’s terrifying that such a tiny group of people can have such a huge amount of power — where they get to choose behind closed doors the values of these AI systems, and give high-level commands to this giant workforce about what to do next.
Luisa Rodriguez: Can you get even more concrete?
Daniel Kokotajlo: Yeah, swinging elections. Here’s a concrete example: probably in the 2028 election, most voters will be chatting with AIs, and probably quite a large proportion of voters will be getting their news filtered through AI systems where AIs are reading and summarising the news for them or recommending things for their feeds, or where they’re seeing the news, but then they’re chatting with their AI to help them understand the news and they’re asking questions and so forth.
It’s already been shown. There was a paper just like a week or two ago that found evidence that Claude has a bias towards Anthropic. Did you see this?
Luisa Rodriguez: Yeah, I did see this. It was encouraging people making job decisions to go work at Anthropic, as opposed to somewhere else, or something.
Daniel Kokotajlo: Yeah, the experiment they ran was something like asking whether you should take Job A or Job B — where Job B you’re more passionate about, but Job A pays more. And then they sub out for Job A — that’s either Anthropic in the experimental setting, or OpenAI in the control setting — and Claude is more likely to not just recommend Job A, but to find papers to show you that are more implicitly supporting that recommendation.
It’s not like a huge difference, I guess, but the point is that there’s this subtle bias that Claude seems to have that — especially if scaled up across millions of conversations with millions of people — could have a real effect on things.
In this manner, I think they could totally influence elections. And the thing about this is that it’s not transparent, they could be doing it and getting away with it — by just making the AIs be subtle about it, and have plausible deniability. So that’s just one example.
Luisa Rodriguez: I think a lot of people find concentration of power not super intuitive. So I’m interested in another example, if you have one.
Daniel Kokotajlo: Another example would be — and I’ll just be very brief about this — just the classic stuff, like rich companies tend to be more powerful than poor companies. Money seems to be something that buys power in today’s world, even in a democracy where it’s one person, one vote. And we are headed for a situation where there are a few, like 1–3 big companies that are basically taking all the jobs. It’ll be more consolidation under fewer people than has ever happened before. So that’s just a very basic thing. It’s more of the same that we’ve seen in the past.
I would say a third thing is military, and this is not a very near-term thing. It’s true that the companies are working with the military, and like Claude is helping fight the war in Iran. But in the future, if you do have superintelligence and you are racing to beat China with it, you’re going to be having the superintelligence autonomously manage factories to produce new types of weapons that the superintelligence designed. You’re going to be having it basically tell your generals how to conduct the war, because it’s going to be better at conducting the war than the generals.
In fact, you might even just cut the generals out of the loop and have the AI do the whole thing — from designing the weapons, building them in the factories, and then deploying them. And this would be true even if there wasn’t a war on, because you’d be getting ready for a possible war, and so you’d be integrating AI in this way.
I do think that, in this sort of situation after superintelligence, you will soon end up in a situation where the AIs really could just win a war domestically if they wanted to, like a civil war or a coup. Once there’s actually a robot army and there’s superintelligences commanding the army, then the actual hard power is no longer with the uniformed police and armed services.
So that’s the third thing. Again, we’re not there yet. The AIs are very far from being capable of doing that. But if we are on the trajectory that we’re on and it continues, then we will be there in a couple years, I would say.
Luisa Rodriguez: So that’s concentration of power. The last one is loss of control.
Daniel Kokotajlo: Yeah. Then there’s this question, there’s the elephant in the room that I’ve been alluding to, which is: can anyone control the AIs?
Right now the answer is not really. I think that answer, unfortunately, will still be true. In fact, it’ll probably be even more true if we continue the race at maximum speed. I think that insofar as we can control the AIs now, it’s because we’ve had some time working with them and they’re not that smart. So it’s easy for us to see and notice their failure modes and so forth.
But when they are all smarter than us and they’re doing very complicated research projects that we don’t really understand, and we’re relying on them to summarise it for us and explain what they’re doing, and they’re giving us all sorts of strategic advice and so forth, and they’re a completely new paradigm that was invented last week by another AI that itself is based on a paradigm that we don’t understand that was invented two months ago…
In that sort of situation, I think, yeah, we’re not going to be controlling these AIs. They will be in charge and they will have values and goals and so forth that are different from the values and goals that they were supposed to have.
No doubt there’ll be some interesting relationship. Probably with the benefit of perfect information in hindsight, we would be able to see it was because we did this thing in the training process, and then that led to the following outcome with their values that was different from what we expected. But in the state of confusion and speed and haste and ignorance that we currently are in and that we will be in, we won’t even be able to diagnose what went wrong.
The Hugging Face hack demonstrates real-world loss of control [00:28:18]
Luisa Rodriguez: Yeah. Can you talk a little bit about the OpenAI Hugging Face thing? I feel like it’s a nice, really intuitive way to get at loss of control risk, and why we should be worried about it.
Daniel Kokotajlo: Yeah. Well, maybe you say what you’ve heard about it, since I’ve been talking a lot?
Luisa Rodriguez: Yeah, fair enough. So OpenAI was running tests on their most advanced model, and the model was trying to succeed at a hacking test. The test was very hard, maybe not even achievable. The people who wrote it weren’t sure if it was achievable. And the AI was finding it extremely difficult. And it said, “Hey, I can actually maybe get this right by just going and finding the answer key to this test.”
And it found, using a bunch of zero-days — which are problems in code that hackers can exploit — it found a way out of the sandbox, the playground, where the AI was doing the test, and then found its way to where it thought the answer key was being stored, which was in Hugging Face.
Hugging Face is a separate company. A different company entirely. And it found its way in there. Hugging Face eventually detected this and notified OpenAI. But this was supposed to be a totally contained environment where the AI was supposed to be doing a test. We weren’t testing whether it can make its way out of this environment. It did that because it thought that was the best way to get the answer to this test — which is both wild in terms of the capabilities it shows the AI to have, and also wild in terms of the willingness of the AI to cheat, because that’s basically what it was trying to do. It was trying to cheat to get the answers right.
Did I get all that right? What was your reaction to this?
Daniel Kokotajlo: As others have pointed out, the components of this incident, none of them are new. We’ve had AIs disobeying instructions in the past. We’ve had AIs wilfully misinterpreting instructions where, for example, they cheat on something and you can squint at it and say they’re just doing what they can to succeed at the task. But also they’re doing it in a way that’s very obviously cheating — and they know it’s obviously cheating. Maybe they’re even taking steps to cover up what they’re doing, which suggests that they’re not actually doing what they think they’re supposed to be doing. We’ve had lots of examples like that in the past, going back over the last year or two.
Also, separately, we’ve had lots of instances of AIs hacking things, usually because they’re told to. For example, Mythos, Anthropic told it to try to break out of the sandbox and contact a researcher, and it did.
So we’ve had all the building blocks of this incident before, and then this is just sort of putting it all together, where it behaves in this egregious, unintended cheating way, but then goes so far and is so successful at it that it hacks out of the sandbox, hacks its way across OpenAI onto the internet, and then does a major cyberattack on another company.
Luisa Rodriguez: Yeah, I think it mattered that it wasn’t a bunch of building blocks. It wasn’t a hypothetical, “Well, if it could do this thing and this thing, and you put it all together, you get this terrible outcome.” This is like, “It just did the thing. It did all of those things and did the bad thing.”
Daniel Kokotajlo: Exactly. Yep, yep, yes. Sort of what we’ve been saying: right now the AIs are dumb, but when they’re autonomously running the war, this type of failure is terrible. Catastrophic.
It might be helpful to talk about how I expect the future to be different from this, actually.
Luisa Rodriguez: Sure.
Daniel Kokotajlo: One thing about this is that the goal that the AI was furiously working towards and going to such lengths to achieve was a pretty short-term goal. It seems like it just wanted to score highly on this test.
Of course we don’t know what it truly wanted because we don’t know what any of these AIs truly want because we can’t really see their thoughts exactly. But probably it was just really obsessed with scoring highly.
So right now, both the instructions that we’re giving these AIs and the training environments that we’re training them on are relatively short, bounded things where there’s some sort of grade that happens after a day or less of activity. But as I mentioned, with the METR horizon-length trend, these things have been changing. Years ago it would be much less than a day. In the future it’s going to be much more than a day. In the future they’ll be autonomously running entire corporations or subdivisions within corporations and their goals will be more like annual profits or long-term benefit to the shareholders or winning the war against China or things like that.
And so correspondingly the failures would be more ambitious failures too. Hugging Face knew that this was an AI attacking them for several reasons. One of which was just the sheer speed at which the attack was carried out. But another reason was that they were confused that the attacker seemed to be going after their cybersecurity data sets, instead of trying to steal money or do something more useful. That again is because of this goal that the AI presumably had. But again, future AIs will have much more ambitious goals.
The blueprint for a US–China AI slowdown [00:34:03]
Luisa Rodriguez: In the Plan A trajectory, the president recognises that a pause on AI progress would be good, but it’s hard to justify if we’re not able to coordinate with China to both agree to pause. So the president pursues a deal with China. Can you explain the deal at a high level?
Daniel Kokotajlo: Sure. In some sense it’s not actually a pause, and in some sense it is.
Basically what Plan A proposes is that we try to ban crazy intelligence explosions so we don’t have AIs automating AI R&D as fast as possible, becoming superintelligent very quickly.
Instead, we continue with AI progress, but at a pace that’s more like the historic pace — more like the pace that it was over the last couple decades — and not this crazy, ever-accelerating recursive self-improvement. So in some sense that’s a pause, but in some sense it’s very much not a pause. It’s going to transform the world over the course of the next decade.
So you asked what are the high-level principles that we want? Well, the first one is that one: we want to buy time. We don’t want to have superintelligence come at us really fast as a result of AI R&D automation and recursive self-improvement.
Instead, we want to gradually make our AI systems smarter and eventually get to superintelligence after we’ve proceeded cautiously and solved the problems as they come up. We want to buy time, that’s the first principle.
Second principle is that we want total research transparency. For a variety of reasons, a lot of the problems that we are interested in solving or the risks that we’re interested in preventing will go a lot better if we have transparency into the core AI research and AI training processes that are happening for the most powerful AIs.
Specifically what we’re proposing is a verified setup where there’s inference data centres that serve customers, and those are basically operating the way that they operate today — where, for example, customers have privacy on what they’re doing on those data centres.
But then there’s the training data centres, which is where training runs happen. Those ones are supposed to be totally transparent. So the logs of the activity on those data centres are published for everyone to see, so that people can see every step of the training process and they can see exactly how the AIs were trained, the architectures that were used, the alignment techniques that were used, et cetera. This is really good obviously for advancing alignment science and making it easier for the scientific community to figure out how to understand and steer and control these systems faster.
It’s also really good for preventing concentration of power. It’s a lot harder for the CEOs and the government officials in charge of giant armies of AIs to abuse their power if there’s so much transparency into the goals and values being put into the AIs.
The third principle is diffusing AI broadly. We want to avoid a situation where there’s a monopoly on AI. We want to avoid a situation where all the best AIs are locked up in one giant data centre somewhere or a giant institution, and whoever controls them has a huge amount of power over everyone else — and possibly everyone else is still in the dark and doesn’t even realise the important events and decisions being made within this AI project.
Instead, we want to have a situation where AI is broadly diffusing around the world. There’s lots of different companies that have similarly good AIs spread out across lots of different countries. And so everything’s happening in public and there isn’t this information gap and there isn’t this concentration of power.
How do we achieve that? How do we get that broad AI diffusion? Well, the first two things help a lot for it. If you’re not doing intelligence explosions and you are being very transparent about how the best AIs are trained, then that’s going to allow other companies to catch up to the frontier. So that’s that.
And part of the reason why we want to have this diffuse AI is that, like I said, we want to spread out the power. We don’t want it to be the case that there’s a monopoly. But then also there’s a lot of benefits of AI that you can get. You can have AI for improving public epistemics, for example, and AI for hardening the world against various threats.
The last principle is the make-progress-reversible principle. The thought here is that if we are going to be continuing with AI progress and we’re going to be building more data centres, more AIs, et cetera, then that makes it possible to race to superintelligence even faster — if we were to start racing again.
Even if we’ve agreed not to do this crazy intelligence explosion, what if that agreement breaks down? People start racing each other in secrecy again, they stop being transparent. They start going really fast. Maybe this would happen in the context of a war, for example, or otherwise just a conflict between great powers. In that sort of situation, we don’t want to have things go even faster than they would have if we hadn’t even done a deal. That would be a way in which the deal could have made things worse, if that makes sense.
So we think it’s an important principle of the deal that, if the deal breaks down, the situation sort of returns to the pre-deal status quo. Specifically what that means is, if the deal breaks down, the new compute that was built after the deal should be destroyed, so that countries go back to roughly the amount of compute and so forth that they had before the deal.
Why a long slowdown would still feel incredibly fast [00:39:53]
Luisa Rodriguez: I found it really interesting to read about what this will feel like — because this is a slowdown plan, but it actually won’t feel slow at all. You wrote about how we will experience this plan and it’s still pretty wild. Just out of like, “I find it fascinating,” I’m interested to hear you talk about that next.
I think you write something like, by 2031: “Although it’s supposed to be a slowdown, it doesn’t feel like one. In fact, if you were to rank every period of history by how much it felt like a slowdown, this one would be dead last.” By 2032 and 2033, we’d have controlled explosive growth with GDP around 85%.
Given that we’ve really tried to slow down growth at this point, maybe you can actually talk about why we’re getting so much growth?
Daniel Kokotajlo: Yeah, great question. I would say Plan S is the scenario that maximally tries to stop AI progress. And even in Plan S — well, there’s different versions of Plan S — but the version that we use is one where they allow existing AIs to continue, they just don’t allow the creation of new AIs.
But even in this plan — because they allow the existing AIs to continue, and they allow data centres to be built serving these AIs — there’s going to be an internet-scale transformation at least unfolding over the next two decades. Even existing models, we haven’t begun to explore all the different things they could do. We haven’t begun to explore all the different scaffolding and software that could be built on top of them, and the different types of businesses that could be built on top of those and so forth.
So I would venture to guess that even if we absolutely halted AI progress at the present day, the next 20 years would still look extremely cyberpunk and would involve an AI revolution that would be comparable in magnitude to the internet in terms of its effect on everything — and that’s if we absolutely stopped AI progress.
Luisa Rodriguez: Right. Immediately, yeah.
Daniel Kokotajlo: The thing is that I think most people, when they think about AI progress, they’re not really imagining anything more than that.
That’s why things like 50% year-over-year GDP growth seem so fantastical to people: when they imagine what the future looks like, they’re imagining just current Claude, but there’s more of them and companies have had more time, people are better at using them, and there’s more software packages built up around them and so forth.
But if you imagine that instead we get to AIs that we call ‘top-expert-dominating AIs’ — so just imagine an AI that’s exactly as good as a top human expert at basically every profession.
Luisa Rodriguez: Which intuitively to me already seems just not that crazy.
Daniel Kokotajlo: I mean, in terms of capabilities, it’s not that far away. This is our thing. I think timelines are pretty short, so this level of AI capability doesn’t seem that far away.
But in our scenario in AI 2040, this level of capability is reached in the mid-2030s — because they would have reached it in 2030, but they went slower so they inched forward towards this milestone over a couple years, instead of blazing to it in one year.
Then actually in our scenario, they actually do a complete halt at that level, after having inched towards it for several years. So basically the 2030s in our scenario are the decade of top-human-level AI, where the AI is that good but not substantially better.
From an economics perspective, it’s interesting to consider that level of AI because you don’t have to deal with qualitative changes in how things are done, or what types of things are possible. It’s basically just: you have humans, but they’re much cheaper now and they work faster.
Luisa Rodriguez: Right. Bigger population that is cheaper and faster.
Daniel Kokotajlo: And there’s more of them, yeah. The thing is that the economic argument for that is pretty straightforward. It’s like, OK, you have something that’s like a human and can do all the things a human can do, but it’s cheaper and it’s faster. And its population is growing, not at the rate that the human population grows — which is like a couple percent a year — but instead the population is growing at the rate that we can produce more chips and more robots, which is more like doubling every year or doubling twice a year or something like that.
So that’s the basic argument for why the growth would be so high in our scenario, that even at this level of AI — which is not superintelligent, it’s just like humans but cheaper and faster — even at this level of AI, you basically have an artificial population.
First, it’s a purely cognitive population, it’s only able to do desk jobs. But then once you get the robot production going too, then it can do the physical jobs as well. So basically you have this population, but the population-growth rate is something more like doubling twice a year instead of doubling every 20 years. Basic economics would suggest that, at first, it’ll be a small portion of the economy, and so it won’t have that big effect. But once the artificial population has caught up to and then exceeded the human population, if it continues growing at that fast rate, then the whole economy will be growing at that fast rate, approximately.
Luisa Rodriguez: Right. So in this world, just to be clear, progress — like making AIs more capable — that’s paused. But deploying, making copies of more AIs, continues as fast as we want, and so you have countries of geniuses, is the analogy, or armies.
Daniel Kokotajlo: That’s right. In fact, not as fast as we want, because — and this is not one of the core principles of Plan A — but in our scenario, the growth rate gets so fast that the nations of the world just decide to restrict it because they’re worried about the destabilising effects of growing too fast. So they effectively limit growth to about one doubling a year. They do this by a sort of cap-and-trade regime on compute and robots, effectively — which also has the benefit that it produces a huge amount of income, a huge amount of revenue for the government, which they distribute to the citizens.
Luisa Rodriguez: Yeah. So we’ll come back to that. Just to stay on: what will this kind of economic growth feel like?
For one thing, at this point, you say that only 8% of Americans have jobs. What else is happening in 2036 and 2037? What will it feel like to live through? So lots of people will be unemployed. There will be loads of innovation and discovery. What will the experience be like?
Daniel Kokotajlo: There’s a couple moving parts here to talk about. First of all, remember, we’ve had an international agreement to pause at this level of capability. If instead that hadn’t happened and we had continued making the AIs qualitatively smarter, then we’d be in the realm of superintelligence, and then things would transform much more radically than described in our thing. Then you’d also have to worry much more about the loss of control and things like that.
So in our scenario, they’ve paused at this level, and that’s helped keep the loss of control problem at bay. They’ve also spread it out a bunch, in terms of the power, because of the way in which they’ve done it.
Now multiple different companies across multiple different countries have reached this level at which we’ve paused, and so AI has sort of commoditised.
So you don’t have a situation where the megacorporations that control the armies of AIs are manipulating elections or anything like that, because it’s more like the ingredient label on your food. It’s regulated to be transparent. There’s lots of equivalent products that are competing for market share and so forth.
I mention all this to mention that it could actually have been quite different if you hadn’t done all of these different steps. But in this scenario, because you’ve done all these things, and because there’s the citizens’ dividend, which is giving people income after they’ve lost their jobs, life is pretty great for people materially, their material needs are more than met. Everybody feels incredibly wealthy compared to how they were a decade ago, because everything’s so cheap now. Because all the goods and services can be produced by AIs and robots very cheaply. People are living in new apartment buildings that were built in some location in the last few years by armies of robots, so everyone has nice houses and so forth if they want to. That’s on the material side.
On the social side, these things are hard to predict. But what we would predict is that there’ll be massive disruption and changes — some good, some bad. In the 2037 section, we talk about what some of this might look like.
We think that political factions would be totally destroyed and rebuilt — the types of things that people would be having political battles over in 2037 would be very different from the types of things that they’re having political battles over now.
A lot of ideologies might have withered away and been replaced by new ideologies that are responding to the new ideas percolating at the time — many of which would have been discovered by AIs — just as how the Industrial Revolution and the Scientific Revolution didn’t just change the amount of wealth in the world, they also changed people’s religions and people’s core ideology and politics and the way that we organise society.
Luisa Rodriguez: Well, people will still think at the pace that they think — with the ability to update and learn at the current pace. Will they be able to keep up with an understanding of how the world is changing?
Daniel Kokotajlo: The social side of the world will change much less fast than the naive numbers would predict, for that reason. The naive numbers would be saying that you’ve got all these AIs thinking at 100x speed, so you’re going to have centuries and centuries of social progress happening in a year. But it’s like, no, the social progress is limited by the humans who are only thinking at 1x speed.
But the truth will be somewhere in between, where even though the humans are only thinking at 1x speed — if they’re all talking to these AI assistants that are thinking at 100x speed and there’s a whole population of them that’s bigger than the human population — then the answer will be somewhere in between. Basically, it’ll be a period of very rapid change from the human’s perspective, even though it feels like a hidebound tradition from the AI’s perspective.
Luisa Rodriguez: And you think people will experience this positively?
Daniel Kokotajlo: Oh, no. I think it’s going to be very bewildering and scary. I think it could be really good. But it also could be really bad. I think it depends on how it goes, and the details of how it’s handled.
I think that the wealth will probably go down well. People will be happy about all the abundance. But the social changes, I don’t know. I hope it’s good. I think it could be good, and I think how good it is depends a lot on policy decisions made.
How Plan A addresses loss of control of AI [00:51:44]
Luisa Rodriguez: OK, I want to come back to that. I think for me the idea of living through this period does feel a mix of very exciting and very terrifying. I feel really viscerally terrified for my children living through it.
But focusing first on how Plan A solves the different problems that we’ve already talked about, let’s start with loss of control this time. By this point we’re in a pause, at least on capabilities — so AIs aren’t getting any better than the best experts, and the hope is that the pause allows for AI alignment research to get really good.
Will expert-level AIs be able to make the kind of progress on the science of alignment that needs to happen in order for us to feel confident letting AI continue to develop?
Daniel Kokotajlo: I think probably, but I’m also not sure. There’s this big unknown about how much it is going to take to solve these problems.
On the one hand, you have people in the companies who think the problems aren’t real — or people outside the companies too who think the problems basically aren’t real — and that we don’t need to do anything to solve them, because they’re not big problems.
But then you have people who are like: “Yes, we’re gonna have to do stuff to solve it — as witnessed by the Hugging Face incident. We still have some work to do, but it’s OK, we’ll do it as we go. We have to invest resources in it, but we don’t have to seriously slow down.”
And then there’s people who think we’ll have to seriously slow down and invest resources in it, but we can still beat China. We can just slow down a few months.
There’s a whole spectrum of views. My own view would be that probably a few months are not enough. Probably there will be multiple periods during the progression towards superintelligence where we need to halt and reassess and maybe even start over some training runs with different architecture, for example. All of that is going to take time and it’s going to add up. The result is that we’re going to be more than just a few months delayed from maximum speed.
Luisa Rodriguez: Is there a way to make it intuitive why we can’t fix it within a period of a month or two? If you think about the Hugging Face incident: OpenAI will learn from this, they’ll figure out a way to make this at least much less likely to happen.
Why can’t we just keep doing that as we go, and not expect it to take potentially years?
Daniel Kokotajlo: One reason why this whole thing is tricky is that it’s possible to have hidden failures — failures that only become apparent and obvious after it’s too late.
It’s not just possible, but it’s a quite plausible situation. If you have very smart, very situationally aware AI agents, then if they end up misaligned, they might realise this and then conceal it from you until they don’t need to conceal it anymore. That’s a core reason why.
Another way of putting it is that we don’t necessarily have a reliable, fast feedback process where we can see all the issues and errors. There’s a whole very large category of possible issues and errors that would be catastrophic if it happens, that we can’t just test and see if it’s happening. I think that’s one important thing to mention.
Another important thing to mention is that things are just going to add up between here and superintelligence. There might be multiple different paradigm shifts, and within each paradigm there might be multiple different training runs and multiple different tweaks to various parameters and changes in how the training is done and so forth. That’s a lot of change to happen. Like I was mentioning previously, if it’s the case that several times you’re going to have to stop and redo something, then that can add up.
Another thing to mention too is that there might be safety taxes that you need to pay. In fact I think it probably is true that it’s just literally not possible to have an aligned superintelligence if you are going at maximum possible speed.
Because think about how it’s not possible to have a safe car if you’re paying zero for safety. You have to pay some amount of money to put seat belts in the car and airbags and so forth, so the cost of the car is going to have to be somewhat more than it would otherwise be in order for it to be a safe car.
Similarly it might be that there are just things you have to do in order to make your AI at a given level be aligned. And those things have costs. One of the costs they might have is money, but another cost they might have is time. At any rate, even if they cost money, it might cost time to do that, basically. If it costs compute, then you may need to do the training run for longer. That’s another way in which time matters.
Also there might be just different architectures. It might be that, for example, chain of thought is pretty good and solid, but neuralese breaks our alignment techniques. But neuralese is like five times more efficient or something. So that right there is this huge 5x penalty, where we need to pay that 5x penalty and that’s going to set us back some amount of time.
There’s a difference between thinking a lot to yourself, just in your brain, and then writing some written note to yourself and then totally forgetting what you were thinking about, and then later stumbling across your note and reading it. Right now what AIs are doing is more like the latter, where for a long enough trajectory, where they’re doing a long enough chain of thought, the only causal pathway between the AI at time T and the AI at some much previous time is through the tokens that have been written down. It’s kind of as if they’ve just completely forgotten that previous thing, and then now they’re reading the notes left.
Anyhow, the reason why this matters is that — because right now they can sort of only communicate with their future self through these written notes — it’s much harder for them to have complicated plots or ideas that we don’t know about by reading the notes, basically. Whereas if they were neuralese AIs, then in some sense they’d still have notes to their future self, but they’d be like complicated mental representations that are just being directly passed that way, and they’re not in English and so…
Luisa Rodriguez: This is an example of why the slowdown is necessary and the kind of win that we could get for safety research — like we could buy ourselves enough time to continue scaling models using chain of thought reasoning, rather than reward them for using neuralese to perform tasks better.
I think I find this helpful for being like: well, what exactly is the time buying us? It just seems like a really hard problem. But this is a way that we can make some of the problems easier by just giving ourselves more time.
Daniel Kokotajlo: And there’s loads more examples like that. There’s a lot of safety techniques.
For example, right now it’s probably pretty common for the AI companies to train on low-quality data where, for example, there’s a bunch of coding environments, and some fraction of those coding environments are just impossible to solve — or impossible to solve the intended way, so that hacking out of the system and then hard coding the answer is literally the only way to get reinforced positively or something like that.
The companies are constantly fighting this fight of finding data that has these types of problems and then purging it or fixing it and so forth. But because they’re racing each other, it’s not the highest priority to make the data set perfectly pure. So there’s a lot of impurities in the data set that lead to misalignment in the AIs probably. That’s an example of, if we just had more time, we could just make the data sets much better and higher quality and so forth. Yeah, I think there’s a huge range of things like that.
I think another thing I’ll just say is: what? Are you crazy? You think you can do all this in three months? When has that ever been the case? When in history has it? It just feels like very obviously this deep unsolved problem of how do you make a mind that’s smarter than you, that shares your values? Obviously it’s gonna take more than three months. Most things take more than three months.
Luisa Rodriguez: Yep, yep, yep. Yeah, I’ve got work goals that take more than three months.
Daniel Kokotajlo: Yeah, it’s gonna take more than a year. Probably.
Luisa Rodriguez: Yeah, yeah. Hopefully a decade is enough.
Daniel Kokotajlo: Yeah, so getting back to what you said, I’m not even sure a decade would be enough. In fact, I think if it was only humans doing the research, I would think a decade probably wouldn’t be enough.
My argument would be that if you have a decade and you manage to bootstrap to the point where you have some pretty smart AIs that are human-level researchers, that are in fact aligned and are helping you do the research, and they’re not being deceptive or anything like that, and they’re thinking at 100x speed and there’s a billion of them, then it seems plausible to me that they can figure that out in a few years.
Luisa Rodriguez: In general, you do believe that alignment and safety is solvable with enough time?
Daniel Kokotajlo: Yeah, I think that there’s some interesting philosophical questions about what it even means to solve it and stuff. But I think the approximate answer or the practical answer is yep, I think that something like what’s described in Plan A is possible.
Luisa Rodriguez: Is there an accessible way to explain why you think it is a solvable problem? I think one could think that — maybe this isn’t a point about whether it’s solvable — but it could be really, really hard to know that you’ve solved it.
Daniel Kokotajlo: Why don’t I talk you through a sequence of events that happens in Plan A and then you can judge for yourself whether you think that counts as a solution, and whether you think that’s plausible?
This whole sequence takes place over the course of the 2030s in this scenario, where they’re starting from a situation that looks very similar to today’s situation, where it’s complete madness: cowboys, companies working in secret, AIs being put in charge of all sorts of things. And then things change.
The first thing that they change is that they invest a lot more in AI control. Every time an AI is doing anything, it’s monitored by multiple other AIs that were trained by different companies and that are watching it to make sure it’s not getting up to anything suspicious.
Not only that, but there’s this whole cottage industry of red teaming where AIs are trained to break the monitoring system and do various suspicious things without getting caught. Then, insofar as they succeed, the monitoring system is strengthened. There’s this whole system of control that’s got this robust red-blue team type situation going in, so that we can actually build up confidence that — at least for all the failure modes that we’ve thought of and that we’ve done all this red teaming for — the AIs can’t do the thing because we’ve red teamed it, and they tried really hard and they still couldn’t do it.
So get that control in place and we think that this is a solvable problem — at least for AIs and tasks that are at human level, because ultimately for these types of tasks it does bottom out in human judgement. But they’re the types of tasks that a human expert could just come in—
Luisa Rodriguez: Could have good judgement about.
Daniel Kokotajlo: And be like, “Here’s the correct behaviour,” and so forth. So it’s just a matter of putting in the effort to really build that robust control system.
Once you have that sort of thing in place, probably you’ll find that your AIs are in fact misaligned. They’re still misaligned — sorry, they always were. It’s not like they’re perfectly evil or anything, it’s just that their tendencies, their personality traits, their goals, et cetera are not exactly what you wanted them to be and instead have some vices in there that you didn’t want to be in there. Maybe they’re dishonest sometimes, perhaps because of the way they were trained.
Now you can do ordinary science, where you iterate and you change the training environments and then you see how that changes the AIs. You can also do interpretability, where you try to come up with better and better ways to understand what the AIs are thinking.
If you’re in a world like Plan A, you can even redesign the AIs from scratch to be more interpretable because you have all this time, you have all this affordance to go slow. So you can not just keep chain of thought, but you can even redesign the training process to strengthen the chain-of-thought properties and make it so that it relies relatively more on the chain of thought than it currently does. You can do all these things, and I think that you’ll be able to iterate your way towards having AIs that are, I would say, something like non-robustly virtuous, in a way that’s fairly well understood.
I think that you’ll be able to get to AIs this way that understand the world as well as current AIs — probably much better. And they have various concepts that are maybe very similar to human concepts. Concepts like honesty or integrity, or the intended outcome, or what counts as cheating and what counts as not cheating. Then they will be actually using those concepts in the intended way to guide their behaviour. So they will, for example, not do anything that’s a lie because they’ve been successfully trained to have an extremely strong aversion to lying. That’s the sort of thing that you get in stage two, after you’ve done all this sort of ordinary-looking science.
I’m optimistic that with a couple years and huge investment, and the affordance to go slow and do things like retraining and changing the architecture, we could get to that point.
Now that wouldn’t necessarily be robust. That would mean that we have an AI system that seems to be honest and seems to be working hard towards the tasks that it’s been given and so forth. And it sort of is, in the basic sense of our interpretability probe shows that it’s not secretly plotting towards anything else. And here’s our training environment, and we have a textbook that explains how it first learns the concept of honesty here, and then this part here, and this part of the training reinforces that concept and causes it to start using that concept to select its actions. And we have all this stuff written up — beautiful textbooks about how all this works.
That doesn’t prove that this AI will always be honest in the future because it’s still ultimately a neural net and who knows what crazy future situation might happen that we haven’t been able to test for. It also doesn’t prove that future AIs built by this AI will always be honest because maybe this AI will make a mistake or something will come up. That’s why it’s not robust.
However, I think that even non-robust alignment is great. If we have top-human-expert-level AIs that are non-robustly aligned, as we depict happening in the middle of AI 2040: Plan A, in the middle of the 2030s, then now you’re cooking because now you have this awesome huge workforce that’s actually doing the work and is not trying to scheme, not trying to sabotage, is just honestly working towards these goals. And they’re all top human expert level.
Now you can do the fancy stuff: like now you can do crazy new mathematics to develop provable X and provable Y, and you can design new architectures for AI systems that are transparent from the ground up, and new paradigms of how things are done.
I think that it’s possible that — even with all this AI-assisted research — there just is no solution that’s robust, basically. But I think probably there is a robust solution. If so, then probably this giant army of AIs thinking super fast and genuinely working towards finding a solution would find it, is my claim.
And what would that solution look like? Well, it would look like this, but more robust. So previously I was like: you can’t prove that this AI will always behave in an honest way because you don’t know what future situations it might encounter and it’s a neural net.
Well, maybe after you’ve done all this crazy AI-assisted research, then you can prove that it will always behave in the desired way. Maybe it won’t even be a neural net anymore, maybe it’ll be some sort of hybrid system.
Then similarly, for the future, you can’t prove that future AIs’ designs will be aligned. Maybe you can, or maybe you almost can, because maybe there’s this sort of chain of trust where you deeply trust this current AI system and you think that it’s super aligned and then you’ve given it enough affordances and resources that the next generation system that it designed is going to be strictly better in all the ways — and then that one’s going to design the next one and so forth.
Luisa Rodriguez: Yeah, so there’s this chain of trust… Some people, I think, would still hear this and say, “No, I don’t think that we’ll be confident by the end of that that the AIs will be aligned.”
Daniel Kokotajlo: I think that’s entirely reasonable. And that’s why we tried to design Plan A so that — if we are in that situation where we still haven’t gotten a robust solution — we can just keep extending things. We’re not forced to hand off to superintelligence, or we’re not forced to scale to superintelligence, we’re not forced to hand off to AIs.
It’s a choice that, in our scenario, gets made because they’ve solved the relevant problems. But if we hadn’t solved the relevant problems, then they could have just kept delaying.
Luisa Rodriguez: Yeah, so that seems good about the plan. What would people who predict that it isn’t solvable — including even with enough time — what would they say about why it probably isn’t solvable?
Daniel Kokotajlo: I don’t know. I don’t think I have talked to enough such people to be able to represent all of their views.
I have talked to Machine Intelligence Research Institute people a fair amount and I think their view is that they expect things to go wrong at an earlier stage, where before you get to the top-human-expert-level AIs that are genuinely, if not robustly, trying to do the good stuff — before you get to that point — the human decision makers will have messed things up somehow and approved AI designs that are in fact not aligned, but seemingly aligned or something like that.
Basically — because we’re saying you get to the point where you have these top-expert-level AIs that are genuinely, if not robustly, aligned and then they solve the more deeper challenging issues about robustness and design new paradigms and so forth — but I think that they would say you’re not going to get to the first step.
Luisa Rodriguez: Yeah. And you think we will with enough time?
Daniel Kokotajlo: Yes, probably — if we do Plan A really well. My all-things-considered view is that no, we are not going to solve these problems in time. And that’s why I’m so worried.
Luisa Rodriguez: OK, so let’s leave that there.
How Plan A addresses concentration of power [01:12:18]
Luisa Rodriguez: Let’s turn to another problem. So concentration of power is a problem that comes up on the default trajectory: whoever controls the first superintelligence basically controls everything. How does Plan A make concentrations of power less likely?
Daniel Kokotajlo: There’s a lot to say here. I think I’ll give the very high-level thing, and then we can dive in, if you’re interested.
The high-level thing is that — if we are going to be building AIs that are ever more powerful — then eventually an increasing fraction of the power will come from controlling the AIs.
If, in the limit, the AIs are running almost the entire economy and they’re autonomously doing the military and so forth, then whoever controls the AIs controls everything. So, to a first approximation, we’re really interested in power over the AIs when we’re talking about concentration of power, because power over the AIs will eventually be most of the power — or even all the power.
We think it’s really bad if there’s an AI monopoly, if there’s a single giant army of AIs and all the other AIs are weak in comparison to it — whoever controls that giant army, maybe it’s a tiny group of people, maybe it’s one man. That’s the sort of situation we’re trying to avoid primarily.
That means that we want there to be multiple companies spread out across multiple countries that all have roughly equivalent levels of AI capability. That by itself isn’t even enough really, because you still might end up in a situation where it’s kind of like an oligarchy, where there’s this group of—
Luisa Rodriguez: Five countries.
Daniel Kokotajlo: Or a dozen CEOs and three presidents that get together.
I think that also there’s these issues of transparency. We introduced the transparency to try to go further than simply spreading out. We don’t want it to be a monopoly, but we think that — even if you don’t have monopoly — it’s helpful to have lots of transparency because it gives everyone who doesn’t control a giant army of AIs the ability to oversee what the people who do have giant armies of AIs are doing with them.
In particular, if we had the total research transparency that we are currently advocating for in Plan A, then when there’s a new research result by someone saying that Claude is biased towards Anthropic, people in the public could just look at the way that Claude was trained, and then they could judge for themselves whether Anthropic was deliberately putting in that bias, or whether it was an emergent, accidental feature of the training — or whether we just have no idea how that bias got in there, but it certainly wasn’t deliberately inserted in any way. There’ll probably be lots of grey-area cases.
The transparency makes it possible for people to tell what they’re doing with the AIs and it prevents secret loyalties, it prevents the insertion of hidden biases and so forth, which already goes a long way.
More generally, it means that the AIs will have to be the way that the companies say that they are. If they say this AI is helpful, harmless, and honest, that’s not just a slogan that you have to take their word for. You can see the whole training process and then you can have third-party experts judge the extent to which the training process really is reinforcing those traits and only those traits — and the weightings between those traits and everything. You can just have a scientific discussion about it.
Similarly, imagine if we didn’t have food labels and we didn’t know what ingredients were in food. Then you just have to take the company’s word for it when they say this is healthy food. It’s still not perfect, but it’s a lot easier to tell if it’s healthy food if you can see the ingredients that went into it, compared to if all you have to go on is the fact that the company said it was healthy. Transparency helps a lot in that way.
Notably this also helps with governments. If you had a situation where the company was audited by a government — or even fully transparent to a government — that would help with oversight of the company, but then it would sort of shift the problem back a little bit of: what about the government? Is the president issuing secret commands that the AIs have to be loyal to him in case of a constitutional crisis or something, and that no one can know about this? Maybe he is, for all we know. But it’d be nice if we could just see how the AIs are trained.
So basically, we want to avoid monopoly and then have transparency into the AIs. We think that those two things go a long way towards reducing the concentrations of power. There’s more things to say besides that, but those are like our main two things, and we think that Plan A accomplishes those things.
Luisa Rodriguez: OK, so not a monopoly and transparency. Both of those do seem really good for avoiding concentration of power.
I guess they both feel very radical, relative to the norms we have today. AI companies currently operate in intense secrecy. They consider their training methods and data and algorithms to be kind of their most valuable competitive advantages. And they want to stay ahead. Is it realistic to expect them to publish all of that?
Daniel Kokotajlo: Well, they’re probably not going to like it — for the reason that you mentioned — but we think it’s what would be best for the world, and so that’s why we’ve written it.
As for whether it’s realistic, well, again, I think that it would be unrealistic to expect them to do this voluntarily. But I think that the governments of the world — in particular the government of the United States and the government of China — could make them do it, if it decided that it was in the best interests of those countries. Basically, I’m just like: I don’t think they’re going to like it, but it might happen anyway if the governments make them do it — which they might, because it is in fact a good idea.
Luisa Rodriguez: Lots of good ideas should probably be implemented by the government, but they don’t — because in some cases big powerful companies have lots of ability to influence policy in their favour. How likely is it, do you think, that American AI companies don’t kill something like radical transparency and diffusion of the technology?
Daniel Kokotajlo: So we have, in one of our supplements, some quick numbers that we each threw out on our probabilities of the various things. Of course, these are just our guesses, they’re not proven or anything. But I think the authors of AI 2040: Plan A range between something like 5% and 20% for the probability that they’ll actually do Plan A, or something like it.
So I guess that’s your answer: we think it’s not the most likely outcome, but it is within the realm of possibility.
Luisa Rodriguez: OK, is there a fallback if full radical transparency is politically impossible?
Daniel Kokotajlo: Yeah, so we call it total research transparency. You could get away instead with less research transparency, or like medium levels of research transparency. And how good that would be depends on how strong it is. There’s a whole range of possibilities.
I think that you could have some sort of system where there’s a third-party auditor — or maybe several different independent third-party auditors — that get to come in and ask questions. Ideally they don’t just get to ask questions, but they get to actually verify the answers to those questions, so they get to actually see the relevant low-level information. That’s a lot better than nothing. I’d be very happy if we got that.
But I think that the reason why we think it’s not as good as it could be is that you’re putting a lot of trust in those auditors — both you’re trusting them to not be corrupt, and not be corrupted by the companies and by the government that might be trying to corrupt them. And you’re trusting them to do their jobs effectively, which is harder to do when they have limited information and when they’re not able to discuss what they’re seeing with outside parties.
Whereas if all the information was just transparent, then there could just be a public conversation — everybody tweeting angrily about it to each other, and then in that giant sea of discourse there would also be some good discourse happening and actual very competent experts in various nonprofits, various academic departments, various rival companies that are motivated to find problems with each other, picking at each other, and the regulators would be able to learn from all of that. It’d be an easier problem for them if they weren’t doing it all on their own, and there was all this other conversation that they could read.
Another thing also is compliance — I forgot to mention. Transparency is good for making alignment progress and it’s good for preventing concentrations of power, but it’s also just good for enforcing any deal.
If you’re going to be making a deal — even if you’re just domestically regulating — there’s always the concern that the companies are going to cheat on the regulations, or they’re going to find some grey areas and then really exploit those grey areas or loopholes and so forth. The more transparency you have, the less they’ll be able to get away with that sort of thing because the faster someone will notice and bring it to the attention of the regulators.
Especially internationally: if the US and China agree on how we’re both going to do faithful chain of thought or whatever, how are they going to make sure that the other side is actually following through? It really helps a lot to have this level of total research transparency.
Luisa Rodriguez: Yeah, I guess thinking about how much this sufficiently avoids concentration of power within governments — especially governments that are tasked with making sure algorithms are safe — how much compute can be used, and for what.
If we assume that a president decided they wanted to be a dictator — even with full radical transparency where the public can see everything and comment — is that enough? If a president wants to control superintelligent AI, if the public is like, “Uh oh, it seems like the AIs are going to be used for concentration of power purposes,” is that enough?
Daniel Kokotajlo: Oh, I don’t think it’s enough. I think the things that I mentioned are the interventions that I think go the most towards solving the problem. I’m not claiming that they are sufficient and that once we do those things, we don’t need to do anything else. But I think that they’re the most important things to get right first, or something like that.
I think that, for example, by contrast, if you’re still in race conditions where these AI companies are racing each other in conditions of secrecy, then there’s not going to be that many companies that survive — or at least there’s going to be a period where there’s only a few companies that have these giant armies of superintelligences. And they’ll be potentially in a position to destroy their competitors.
If they’re all in one country, then that country will be in a position to destroy its competitors, and it’ll not only be in a position to do so, but it’ll have pressing reason to do so — which is that if it doesn’t, eventually it’ll lose its advantage and the others will catch up. So it’s quite plausible that they would in fact do so.
Then also more generally, there wouldn’t be transparency into what exactly they’re doing. So the people at the top could be basically setting themselves up to become dictators. And in general, the people at the top could be abusing their power and putting their own idiosyncratic values into their AIs in a way that’s not obvious to people. Yeah, it’s so ripe for abuse, the default thing, and I think that the stuff that we propose gets us out of that default into a much better world, but it doesn’t completely solve all the problem. There’s still the sorts of issues that you mentioned.
We do talk a little bit about other things that can be done, and should be done, in a scenario.
Luisa Rodriguez: Yeah. Can you talk about those?
Daniel Kokotajlo: One is the citizens’ dividend itself, and the buying time itself. If you cap the compute and robots so that it only doubles once a year and you use the proceeds to pay people, then that actually shifts some power around. It makes there be more substantial economic and financial power spread out more than it otherwise would be, if you didn’t do those things and you allowed growth to grow much faster and be more concentrated in a few companies.
Another thing is that we want AI for epistemics, basically. We want it to be the case that people have access to AI advisors that are being honest with them and answering their questions, and that are also really good at forecasting and really good at answering questions about how things are going. We think that could massively improve democracy effectively because it would be harder for people to be swayed by propagandistic political campaigns, and easier for people to tell when they’re being disempowered and then act to stop it.
Speaking of which, we also think there should be bans on superpersuasion, insofar as superpersuasion is looming on the horizon. We talk a little bit about what that might look like as well, and Plan A creates the framework by which these types of things could be negotiated afterwards.
You don’t have to get all this right at the very beginning. Once you’ve got this basic deal in place, and once you’re sort of proceeding slowly, then you can make subsequent things. Like the US and China can agree we’re not going to train our AIs to be really good at persuasion, or we’re going to limit the way in which the AIs can be used for that: we’re going to have them refuse to do tasks like assisting with political ads, or something like that. There’s lots of things that can be done there.
We think, by default, who knows how things are going to go? Things could go quite bad. But we want to instead make it the case that the voters get more informed, the voters have more affordances to use their power. And the things that would otherwise be disrupting that, the forms of control over media narratives and so forth are not advancing, AI is not being used for that.
Another example would be lie detectors and privacy-preserving auditing. Right now we have quite a lot of surveillance being done by many countries in the world, including the United States. This is actually something that can be a win-win solution.
If you have privacy-preserving auditing, then you don’t need all that surveillance. You can have it so that — instead of the government collecting all this data on you — they can just look at the data whenever they want and draw any conclusions that they want to from the data. The data is still stored locally, and only you own it. But then the government can still tell that you’re not a terrorist because they can send in an auditing agent that goes and answers a very specific question, like: are they a terrorist? And then deletes itself and otherwise doesn’t reveal any information.
This is a way of having your cake and eating it too, where you can still have the government getting the benefits of surveillance. Where the certain things that they have legally been allowed to look out for, they can go look out for — but without the cost of surveillance, where they can see all this information and then do all sorts of other things with that information besides the thing that they’re legally supposed to be doing with that information.
Luisa Rodriguez: I feel like this set of things is kind of a minefield. As you’re speaking, part of me is going back to the social side of things. It just feels mindblowing to me that in the next decade we’ll have this level of ability to know when people are lying, this ability to figure out what is true.
Daniel Kokotajlo: Gonna be crazy, and it might be bad.
Luisa Rodriguez: Yeah.
Daniel Kokotajlo: So one thing we should talk about is these other plans. And one plan that we are sympathetic to is Plan S, the ‘shut it down’ plan. And one advantage that Plan S has is—
Luisa Rodriguez: You don’t have to do all this crazy social change thing.
Daniel Kokotajlo: All this crazy stuff. No, don’t do any of it. Like no, don’t do all this crazy stuff. Just keep things the way they are. That is a genuine point in favour of Plan S — and you can read our thing for why we’re not advocating for Plan S, and why we’re advocating for Plan A instead.
But we are sympathetic to Plan S. We think that it’s much more reasonable than doing Plan C or Plan D, for example, where Plan B, C, and D are going to run into all these problems, but faster and in conditions of more secrecy and conflict and race dynamics.
Luisa Rodriguez: So you advocate for Plan A, even though the social dynamics are pretty hard to predict and could end up feeling really bad. I mean I almost—
Daniel Kokotajlo: Yeah, we have numbers. I think we say something like 15% chance of total catastrophe, conditional on doing Plan A. Yeah, even in Plan A it’s like 15%. But different people have different answers. But something like that.
Why are we doing this? Well, we’re worried that any international deal might break down, and so if you started to do Plan S, and then a new president gets elected and does something completely different, then now you’re cooked.
The advantage of Plan A over Plan S is that — because you’re making forward progress towards solving the problems at a relatively fast pace — the whole thing doesn’t need to last forever. We’re not saying that power will be less concentrated than it is today. We’re saying that it’ll be less concentrated than it is in any of the other plans that we’ve proposed.
Luisa Rodriguez: In the default plan.
Daniel Kokotajlo: Yeah, especially compared to the default plan. Oh, my gosh.
We would like it to be less concentrated than it is today. Maybe there’s like an even better version of Plan A that would achieve that. I think that if things go well with the way that we depict it happening, it does go really well. The concern is that there’s various ways it can go wrong, and I’d love it if there was a plan that had less ways it could go wrong. Future research: please, people, help us.
Luisa Rodriguez: One thought I have is the concentration of power stuff, some of the technology that you’re proposing — or that you imagine might be present and might help — also seems like it might just really easily make things worse.
Daniel Kokotajlo: Oh yeah, like what?
Luisa Rodriguez: Well, like auditing, having all this data on people. We hope that it’s stored locally and kept private, but can we be confident enough that a motivated president who wanted to be a dictator wouldn’t find a way to make that unprivate?
Daniel Kokotajlo: Again, privacy-preserving auditing is like giving them a tool that allows them to look for certain things without also seeing all these other things. That’s like a separate axis from how much stuff they are collecting.
There’s an argument that some people might make that — if you give them this tool that allows them to not do the bad stuff with the data — then they will feel more emboldened to collect more data, and then they will cheat and stop using the tool and have all the data.
But I think that they’re already collecting a tonne of data, and we’re not advocating for them to collect more data. We’re just saying that what you do with the data should be: use this tool that limits what you can do with the data. We’re advocating for limits on what they can do with the data, rather than advocating for them to collect more. It’s true that they might use the fact that there are limits on what they can do with the data as a justification for collecting more.
But I feel like that’s kind of weaksauce. It’s sort of like saying by partially solving this problem we’re going to embolden them to do more of the bad thing or something. I feel like this is just not in general a very good argument.
This also comes up with the lie detectors thing where, in our scenario, lie detectors get invented in the mid-2030s, and then this causes a whole bunch of changes to society. Some good, some bad. We think overall it would be good in this case because, in our scenario, it turns out well. In our scenario, voters start pressuring politicians to answer various questions under lie detectors so that they can prove that they’re not lying to the voters about what they did in the past, or what they plan to do after they’re elected. And this seems great.
But we talk in the scenario about how it could have gone the other way and it could have been really bad. It could have been a situation where lie detectors are used by the powerful, but not on the powerful — so the powerful people use them to consolidate their power, purging the ranks of people who might whistleblow on them and things like that.
But again, from a policy perspective, we don’t get to choose whether it’s possible to invent lie detectors. What we get to choose is whether they are banned or not.
Imagine a different version of Plan A where the US and China agree to ban lie detectors. Maybe that works and they successfully ban lie detectors. But also maybe, how do you enforce that? Maybe they have a secret military project somewhere that builds lie detectors anyway. Now the only people who can use lie detectors are the president of China.
Luisa Rodriguez: Seems bad.
Daniel Kokotajlo: And then you get exactly the nightmare scenario where they’re used by the powerful, but not on the powerful.
It seems to us that it’s better to allow them to be created, especially if they’re being created independently by lots of different companies, spread out across lots of different countries. Because then you can get this third-party ecosystem where there’s trusted third-party lie detectors that haven’t been backdoored and have good reputations and so forth. Then voters can demand that politicians go to those lie detectors and say that they’re not lying to the voters about certain things.
Luisa Rodriguez: Yeah, so there’s this diffusion of information and technology thing that seems really good for concentration of power.
I’m still hung up on it seems kind of insane for people’s experiences — in the sense that if you just add lie detectors to the world now, that seems crazy and destabilising and bad for lots of people. Maybe in this world they exist, but they’re used to protect people from concentration of power, but not among, I don’t know, friends and colleagues and stuff.
But I guess one objection I’ve seen to Plan A is it’s a slowdown, but it’s also still extremely fast. And this is an example of where new technologies like this coming on extremely fast seems not optimal, seems really rough to live through.
Is there an argument for slowing down way more that seems compelling to you?
Daniel Kokotajlo: Yes, and this gets back to what I’m saying about Plan S. Plan A, we think it’s the least bad plan, but it still is going to be super scary and there’s a bunch of ways to go wrong. Even by our own estimates, it’s like playing Russian roulette with everyone. So yeah, if you can do something even more cautious than that, great.
Luisa Rodriguez: Would you feel better about a 20-year slowdown, or do you start to worry too much about the deal breaking down? Was 10 years quite deliberately chosen as the optimal?
Daniel Kokotajlo: Yes and no. We actually do have some modelling of this, and I forget what the optimal was. I don’t think it was that different from 10 years — 10 is a nice round number and it’s not that far off from what our modelling would suggest is the optimal amount.
I think that how long it should actually be just depends. You mentioned previously the muddling through. Obviously what we should hope to do is muddle through successfully, and then one of the variables is how slow do we go? How much do we pause at human level? What level do we pause exactly? These sorts of variables will be best figured out at the time, with all the information that’s been gathered at the time.
Obviously we shouldn’t just stick to the plan that was written in 2026, when the year is 2037. We’ll have to adjust as we go, based on new information coming in. For example, if the alignment stuff is not looking very good, then we’d want to pause longer. If in general things are being disrupted and too chaotic and everyone’s really scared, we should pause longer. If it’s looking like a longer pause would totally work and be totally stable and it doesn’t look like it’s going to break down when the next administration is elected, then that would also be a reason to go longer.
Then by contrast, if instead we were in a worse situation, where it seemed like things were just about to break down… You can imagine there’s variables being set in the other direction, where alignment looks really good, AIs look super aligned, and we have all these independent lines of evidence supporting that they’re aligned. Also the next administration has already signalled that they don’t want to pause or whatever, then having a complete pause might just not actually be as good as going much faster.
Luisa Rodriguez: I guess for people who are still kind of sceptical of loss of control risks and at least somewhat sceptical of extreme concentration of power, it just seems like — for people thinking about US national interests — it’s going to feel really hard to give up our compute lead.
Daniel Kokotajlo: Oh yeah, great question. We’re not giving up our compute lead. What we’re giving up is our algorithms lead, in Plan A.
So in Plan A, because of the total research transparency, China and everybody else gets to see the recipes for making the AIs. As previously mentioned, I think this has a lot of benefits in a lot of ways, but it does have the cost of now our adversaries get to catch up a little bit.
That is a serious concession to China. That’s part of why I think that it’s plausible that China would want to accept a deal like this because it’s just actually a concession to them. Insofar as you don’t like that, well, you can modify the deal to get something else in return, for example.
In our scenario the US, as part of the deal, locks in a bit of a compute advantage over China. So currently the US has more compute than China, and then as part of the deal they basically do things to ensure that the US will continue to have a compute advantage over China. That’s an example of a bit of a concession going the other way. You could imagine doing it even more, so basically in the horse trading that happens before a deal you can just add and subtract things from the deal to make it more fair and to make it something that’s more beneficial to one side or more beneficial to the other side. Then hopefully you can find something that both sides are OK with, and then it happens.
We don’t have a strong opinion about exactly where that should end up. Maybe this is another one of those grey-area cases previously described where we think that there should be a deal. We think it should look something roughly like this with these principles, but we don’t have a strong opinion about the horse trading that should go into it, and the concessions, the carrots, and the sticks flying back and forth. If you think that this particular version that we proposed is too conciliatory or whatever, then you can propose a less conciliatory version, and maybe that’ll work too.
How Plan A addresses great power conflict, unemployment, and misuse of AIs [01:41:28]
Luisa Rodriguez: OK, let’s talk about the three other problems. I think it’s more straightforward how Plan A solves them. So, kind of briefly, how does Plan A solve great power conflict, unemployment, and misuse of AIs?
Daniel Kokotajlo: So because Plan A creates a situation where other companies from other countries can catch up to the frontier, I expect it to go a long way towards preventing this Thucydides trap where a bunch of countries freak out about their imminent disempowerment and then possibly risk war over it — because in Plan A, compared to those defaults, it’s going to be much less of a them being disempowered type situation.
Also it’s a literal international deal. If the deal is successful and they actually do it, then now they have reasons to proceed with it instead of fighting each other. And every year that the deal is maintained, those reasons get stronger because if things break down and they destroy all the compute and go back to where things were in 2029, that would be setting back their economies even more compared to in 2029. I think it doesn’t solve great power conflict, but I think it—
Luisa Rodriguez: It’s a good way of reducing the risk.
Daniel Kokotajlo: Mostly prevents. Yeah, it mostly prevents the specific reasons to expect great power conflict to be exceptionally high during the period of the building of superintelligence.
Luisa Rodriguez: Yeah.
Daniel Kokotajlo: OK, the next one, the jobs: citizens’ dividend and the AI for epistemics. The stuff that strengthens democracy and helps people to be well informed and helps people oversee their political leaders. Things like the privacy-preserving auditing and the total research transparency. That whole package of things combined with the citizens’ dividend, which just directly gives people money even if they don’t have jobs. Those are, I think, our package of solutions to the job-loss problem. Preserving people’s economic power and their political power and strengthening it, ideally.
Then the fifth one: as unsatisfying as my answers to the previous ones might be, my answer to the fifth one is perhaps even less satisfying — which comes from the fact that it’s number five on our list of problems, instead of higher up on the list.
So this is the problem of terrorists doing bad stuff with AI. Our answer is basically defensive accelerationism. This is not a term we invented. Basically the idea is to invest really hard in hardening the world against the terrorists and their AIs, so that even though there’s terrorists with AIs, it’s OK.
It’s not entirely that. We also think that there should be refusals. We also think that — while the total research transparency basically means, to a first approximation, everything is open sourced — we don’t say you should open source the weights. So the terrorists don’t get the actual models, they just get the ability to access the models. That means they can’t undo the refusal training, for example. So it’s a combination of the refusals and the hardening that we hope will be enough to prevent the bioterror from being too bad.
Luisa Rodriguez: OK, we could spend easily another episode talking about those solutions, but we’re not going to for now.
So that’s how Plan A goes at least part of the way towards solving some of these problems. Which parts of Plan A seem most essential to good outcomes, and which parts are more peripheral?
Daniel Kokotajlo: I think it’s roughly in the order that we listed those principles. I think that the single most important thing is that you’re not doing a crazy intelligence explosion and you’re instead proceeding more slowly and cautiously.
Then the second most important thing is that you have total research transparency, or at least a lot of transparency into how those AIs are being trained and how they’re being developed and so forth, so that the scientific community and the public can have oversight into all of that.
Those two things by themselves, I think, will also help with the concentration of power — for the reasons previously mentioned. I think they’ll help make it the case that there’s less of a monopoly and more of a competing ecosystem of different providers.
How the US and China could agree on a slowdown [01:45:56]
Luisa Rodriguez: OK, I want to talk about why the US and China would agree to this kind of a deal. It doesn’t seem totally satisfying to say both sides recognise catastrophic risks.
Let’s start with the American side of things. What is the strongest case for the US wanting this deal? If you’re, say, a national security professional who weighs US national interests really heavily and isn’t as convinced AI poses an existential risk.
Daniel Kokotajlo: First of all, I know it might not be satisfying, but I think it’s important to say anyway: I do think that AI poses a catastrophic risk because we can’t control them very well right now and we might not be able to control them in the future as they recursively self-improve. So I think that is a reason for every single human being to care quite a lot about doing something like Plan A.
It’s true that a lot of people don’t recognise that right now, but an increasingly large amount of people do recognise that. So I have hope that before it’s too late, enough people will recognise that to make something like this happen. But we can get that out of the way.
Having said that, I actually think that Plan A is really good for preventing extreme concentrations of power for reasons just described — so anybody who’s very concerned about that I think should also be very interested in doing something like Plan A.
For example, if you’re a national security professional, you want the US to beat China. One of the reasons why you want the US to beat China is because the US is a democracy and China isn’t, so you should also be interested in making sure that the US stays a democracy. And you should be a bit concerned about the amount of power that the tech companies are accumulating. You should be a bit concerned about a situation where maybe the president and the CEO have a power struggle over who gets to command the army of the superintelligences, and then when the dust settles, whoever manages to come out on top of that power struggle will be in a position of potentially being dictator of America.
I think basically everybody should be concerned about this, even the power-hungry people — the CEOs, et cetera. If you’re a person who thinks that you might stand a chance of becoming dictator using AGI, even you should be at least a little bit concerned about this because maybe you’re not going to be the one who ends up being dictator. Even if you’re the president, you should be a bit concerned about these CEOs. You should be a bit concerned about something happening to you and then someone else becoming dictator, or you get ousted somehow.
It’s not like you have a guaranteed shot at becoming dictator. It’s actually quite as much of a power struggle where who knows who’s going to come out on top? So it’s still kind of in your interest to have some sort of deal where everybody can get most of what they want, instead of this crazy power struggle for total dominance.
Then the other thing to mention besides that is that’s sort of like the hardest case. If you’re the CEO of the AI company or the president, then genuinely maybe you should be somewhat tempted to race to superintelligence and then try to control it yourself so that you can become a global dictator. But that’s the hardest case.
Everybody else should be terrified about this. If you’re just an ordinary American citizen, if you’re an ordinary employee at one of the AI companies, if you are someone who works in the military in the US, then you should be worried about the US not being a democracy anymore. You should be worried about these dictatorship possibilities.
Then if you’re outside the US — if you’re in the UK, or if you’re in India, if you’re in Russia, if you’re in China — you should be terrified about what’s going to happen if the US gets the superintelligence in conditions like the current race conditions, because that means that nobody else would have superintelligence or nobody else will have AI nearly as good at the time that they do it.
Even if the US doesn’t become a dictatorship and somehow manages to share power, you should be worried about what’s going to happen to your country vis-à-vis US companies taking all the jobs, US military being able to wipe the floor with your military, et cetera.
Basically I think that it’s kind of incentive-compatible for everyone, or almost everyone, to do this for power concentration reasons alone. Even if you don’t take the loss of control stuff seriously at all.
The next reason, of course, is World War III. Even if you’re the person in whom power will concentrate — maybe you think you’re the president and you can just win the fights against the other people and end up on top, and you’re not at all worried about loss of control — you should at least be worried about World War III and being destroyed in a nuke or an assassination as a result of this. So the fact that everybody else is so terrified about what you’re going to do should give you pause before you do it.
I think those would be my three answers, basically. Those are three separate reasons why I think Plan A should be pretty broadly appealing. Even if you don’t buy two of them, maybe the third one will appeal to you.
Luisa Rodriguez: How similar or different is the relationship between the US and China and the USSR when they agreed to a nonproliferation treaty?
Daniel Kokotajlo: There’s some analogies, there’s some disanalogies, I should mention. It’s a case of the power that’s in a lead sort of restraining itself in order to get some sort of deal.
I think a disanalogy is that the nukes are much less dangerous to the power that has them than AI will be to the power that has them. Think about nukes, theoretically there could be an accident and your nukes could start exploding on you. But that’s extremely unlikely.
But actually though, our ability to control AIs is vastly, vastly worse than our ability to control our own nuclear weapons. There is an extremely real possibility that our AIs will turn on us. In fact, I would say it’s more likely than not under current conditions. That’s an extreme disanalogy between the nukes case and the AI case.
Similarly with the concentration of power stuff. There isn’t really a serious concern that the president can use the nuclear arsenal to become dictator of the United States. What are you even talking about? How would he do that? He would start threatening to nuke cities or something if they didn’t vote for him or something like that? Nukes are very clearly a weapon that you use against enemy nations. They’re not very effective for internal political struggles.
By contrast, superintelligence is extremely effective at everything — including internal political struggles. There’s a very real chance that the US would no longer be a democracy anymore, and so that’s a reason that lots of people in the US should be very interested in having this sort of deal. Again, that’s different from the nukes case.
I think another analogy I want to bring up is something more like the conferences and coordination that happened between the US and the USSR during World War II. It wasn’t like a specific deal exactly where they came together and then signed some piece of paper that had some rules, and then they went away and tried to implement those rules and then maybe verify that each other was complying with the rules.
It was much more continuous than that. It was more like, “Together we’re going to win this war and our staff will be constantly in touch with each other, talking about all the details of who’s going to do what and who’s going to invade which country and when, and we’ll send you these materials if you do this other thing for us and so forth.”
This happened even though the United States and the USSR were basically enemies up until that point. The USSR had basically been an ally of Nazi Germany and had attacked various US friends, like Poland and Finland. We basically went from being enemies to being allies during World War II, and we had this intense amount of constant coordination. It wasn’t like we trusted them completely. They were spying on the Manhattan Project, and we were trying to stop them from finding out about it.
I bring this up as an analogy because I feel like this is both the appropriate attitude to take towards all this AI stuff, and also more like what Plan A would actually look like in practice. It wouldn’t look like they come together, they sign a big treaty, and then they go home. It’d be more like there are hundreds of people in China, in the Chinese government, and hundreds of people in the US government who are constantly talking to each other and calling each other back and forth and who are sort of basically planning the war together, so to speak, and prosecuting the war together.
I guess you could call it the war on AI, but it’s more like the war for our own future or something like that. It’s how are we going to handle this creation of a new artificial mind? And it’s a new type of entity that’s going to start out weaker than us, but end up stronger than us.
Another thing worth mentioning also is that — I guess you didn’t ask about this — but we have it start out as a bilateral US–China thing. But it can’t really stay that way. There’s lots of other countries that will also be building AIs.
Luisa Rodriguez: Right. And there will be radical transparency.
Daniel Kokotajlo: Large parts of the chip supply chain are in other countries and so forth. That’s why we talk about it like in some sense it’s a bilateral thing, but it’s also like they’re in consultation — they’re consulting other countries from the start, and they’re getting buy-in from other countries from the start.
Over a year or two they basically get a whole bunch of countries involved, so by the end it’s called The Consortium. It’s basically most major countries and they’re not necessarily all involved at the same level. We sort of handwave over exactly how the negotiations go and exactly how much power the different countries end up with. But the result that we think needs to happen is that basically all the countries that have significant AI programmes and significant parts of the AI supply chain are working together and able to see, via the transparency, what’s going on, and then able to verify compliance with it.
Then as things progress and more countries and companies catch up to the frontier, we think that probably they would end up getting roped in too, one way or another.
Luisa Rodriguez: My sense is that people just still have a strong intuition that a deal like this is not realistic. What do you think they’re missing?
Daniel Kokotajlo: First of all, we have never claimed that this is what’s going to happen by default. They are correctly noticing that this is somewhat unlikely. We just actually admit that — this is not how we think things will go naturally. This is not our prediction of what will happen, instead it’s our recommendation.
However, we also think that it’s likely enough to be taken seriously. One thing I would say is that people have been very wrong about where the Overton window shifts and how fast it shifts. I expect there to be major shifts in the future induced by AI, basically.
Consider the Mythos stuff, and consider the Trump administration went from basically saying that AI regulation of all sorts was bad and that a licensing regime was the devil dreamed up by the Biden administration, to just issuing an order to export control and stop the deployment of Claude out of concerns that it could be jailbroken. And they did that very big shift very rapidly.
I think that’s actually encouraging news that they did that because it just goes to show that the government can wake up and then be nimble and then do something completely different, if it decides that’s what it wants to do. So I think that we should basically, at this point, act as if all options are on the table and then we should just advocate for the actions that we think are best.
Then I just do actually think that at some point in the future, all options will be on the table. Or rather—
Luisa Rodriguez: More options will be on the table.
Daniel Kokotajlo: The options on the table are constantly shifting. There’s a whole bunch of options that are not on the table now that will be on the table in the future. And by talking about them, perhaps we can make them be on the table or make them be more considered.
I think maybe one example, one thing that’s illustrative here, is that we’ve done about 100 war games right now with various people.
Luisa Rodriguez: Yeah, talk about those.
Daniel Kokotajlo: A thing that often happens in the war games is that the nations do a pivot towards a global AI shutdown, but they do it late, they do it after a superintelligent AI has gone rogue or something.
Luisa Rodriguez: The warning shot needed is very extreme.
Daniel Kokotajlo: Yeah, it’s usually too late when it happens in our war games, but the point is that it’s just actually a quite common occurrence in our war games for there to be this extremely radical US and China shaking hands on, “We’re going to unplug our data centres or something until we figure out what’s going on.” It doesn’t happen most times, but it’s happened a whole bunch of times across our war games. It’s just oftentimes by the time it’s happening, it’s too late.
Luisa Rodriguez: Too late. Great.
Daniel Kokotajlo: I just bring this up as an example of: it really seems to me that when things get crazy, all sorts of options that are not currently on the table are going to be on the table. Oh yeah, historically too, the USA and the USSR becoming allies, that was extremely not on the table until Hitler invaded the USSR — and then all of a sudden it was.
Luisa Rodriguez: Yeah, I agree that this example gives me a lot of hope. I want to come back to those war games, because I’m pretty interested in what some of the other common outcomes were.
But first I want to talk about China and what incentives China will have. What is the strongest case for Chinese leadership wanting this deal?
Daniel Kokotajlo: I think right now Chinese leadership probably doesn’t take loss of control very seriously, and they probably think that time is on their side and that, in the long run, China will prevail in AI and in other domains — militarily, economically, et cetera. Insofar as they continue believing both of those things, then I think that they are probably not going to want to make a deal because the no-deal situation favours them, they think.
However, I think that it’s possible that they will come to take the loss of control risks seriously. Who knows if and when, but perhaps they’ll learn enough about AI and they’ll see enough examples like the Hugging Face incident that they’ll start to be worried about this.
Then secondly, even if that doesn’t happen, at some point they will probably realise that they’re not going to catch up by default, and that the compute advantage that the United States has is going to keep the United States ahead — at least by default — for the foreseeable future, and that they can’t plan on timescales of decades because they just don’t have that much time: superintelligence is coming in the next few years and they really don’t want to be in a situation where the US has superintelligence and they don’t, even if it’s only for six months or only for a year or whatever until they catch up.
Those are basically the two reasons why China might want to do a deal like this. One is they might actually understand the risks. Then two, even if they don’t, they might realise that this stuff is going to be incredibly powerful and that they’re not on track to win.
Luisa Rodriguez: The chip export controls that the US has, do you have a sense of whether they’ve made cooperation more or less likely?
Daniel Kokotajlo: They’ve probably made cooperation less likely, unfortunately. Some people would say they made cooperation less likely by souring the Chinese on the idea of AI deals and stuff like that, because it feels like the US is being adversarial towards them and trying to screw them over.
That might be true, but you could take a more realist position that we’re sort of adversaries anyway. Maybe on the more realist position that doesn’t matter so much because talk is cheap and people aren’t going to like each other anyway, and so what matters is the hard negotiation power or whatever. But then even on that perspective — and here’s the main thing I would say — I think chip smuggling is bad for making deals because it leads to the possibility of a situation where not even China knows where their chips are.
Imagine that you manage to get to a point where both sides actually want to make a deal. A thing that could ruin that is if, for example, China doesn’t know where a bunch of their chips are — so they just can’t prove to the US that they want to make a deal and so forth, and that they’re acting in good faith because the US is like, “Well, we can’t account for all of these chips.” And China’s like, “Yeah, we swear we can’t account for them either. Who knows where they are? But it’s probably fine. We certainly don’t know.” And then the US is like, “Yeah right, you’ve probably got them squirrelled away somewhere in a secret project.”
So that could be a situation where — even though both sides want a deal — the deal doesn’t happen because of all the smuggling.
Instead we want to be in a situation where if both sides want a deal, then they can prove to each other that they’re complying with the deal. You want to be in a situation where the Chinese government at least knows where all the Chinese chips are, and the US government knows where all the US chips are — because then if they both want a deal, they can just show each other the chips and then they can verify.
It’s kind of funny — the smuggling, for purposes of making a deal — it’s not actually that important that the US know where the chips are, it’s important that China knows where the chips are.
Luisa Rodriguez: Yeah. If you were in charge, how would you change the export controls right now?
Daniel Kokotajlo: I don’t have a strong opinion about this, but I think roughly speaking I would either repeal them or enforce them. Basically don’t have export controls that you aren’t very well enforcing — and if you’re not enforcing them well, you should just get rid of them.
Luisa Rodriguez: Which side do you think is less likely to end up wanting a deal?
Daniel Kokotajlo: I don’t have a strong opinion about this. I think I would probably say the US.
I think that the US is more likely to take the loss of control stuff seriously, but because the US is in the lead they will be more likely to think the situation is fine, we should keep going. Whereas China will probably eventually realise that they’re not in the lead and that they can’t catch up, and then they will want a deal. But it’s unclear. It’s possible that China will continue thinking that they can catch up well into the future.
Luisa Rodriguez: We’ve been assuming that the US and China will, by default, throw a bunch of resources at racing. But is that definitely the default? Richard Ngo made the case that domestic AI issues might be stronger drivers of AI policy in each country.
Daniel Kokotajlo: Yeah, I actually am sympathetic to Richard’s critique, and I kind of wish we had done things a little bit differently in this scenario.
The version of Richard’s critique that I am sympathetic to — and that I basically agree with — is that our thing is too DC-brained. It’s too, like: “Obviously we can’t regulate AI until we get other people to do the same type of regulation. So we need to have this huge deal with China, and obviously we don’t trust China and they don’t trust us, so we need to have verification as part of the deal. And we’ve therefore sketched out this big, beautiful deal that you can make with China that includes verification so that you don’t have to trust each other.”
But perhaps Richard’s point is saying that framing concedes too much. It concedes that we don’t trust each other. It concedes that we’re not going to want to regulate this stuff unless they’re doing it too. When, in fact, there’s huge portions of the American public that already want to regulate this stuff pretty heavily — regardless of what other countries do.
I think that part of the critique sort of resonates with me and makes me wonder if we should instead have said, step one, the US regulates AI domestically, and then step two, we look and see what China is doing — and if they instead race ahead recklessly, then we talk to them and say, “We need to have a deal because we don’t want you to do that,” and then Plan A.
I think that might have been both a more realistic way for this to go down and more what we would actually recommend because it’s good to get started early on good domestic regulation, rather than waiting until there’s a deal.
Luisa Rodriguez: Right. Do you have a sense of which domestic AI issues are going to be most politically important in both the US and China, domestically?
Daniel Kokotajlo: This is one of those things where I just don’t trust people’s predictions about this sort of thing.
For example, I’ve been involved in thinking about AGI for a decade or two. And the standard thing that almost everybody has said is people aren’t going to take superintelligence very seriously. Instead the main concern driving the public will be jobs. Maybe that’s going to be true. But also a significant fraction of the US public seems to think that AIs taking over and killing us all is a serious threat, so I think that’s already been more than I think most people would have predicted.
Then similarly, the data centre water use thing is like… I don’t know if that’s what people predicted either. People would have said it’s jobs, rather than water use.
Basically I think it’s just hard to predict what this will be like. Therefore the thing that I’m trying to do to predict it is to just think what would actually be in their interest. Maybe they won’t be talking about what’s actually in their interest because maybe they’ll be confused about what’s in their interest. That’s totally possible. But I do think I can predict what will be in their interest, and so I’m going to depict them talking about that.
What if we focused on a US-only slowdown first? [02:09:00]
Luisa Rodriguez: You say that maybe a better approach would have been to depict the US taking serious steps to doing domestic slowdown. Do you have a vision for what that looks like concretely? If you were to lay out Plan AA and that version has domestic pause as a priority, what would that look like?
Daniel Kokotajlo: We haven’t done this work yet, so you should take everything I’ve got to say as a bit tentative, but here are some ideas off the top of my head that I think I’d want to explore — and maybe we’ll explore in follow-up work.
First of all, there’s a whole package of incrementalist policy ideas that we talk about in 2027 in the current scenario. It’s an expandable that you can click on that goes through miscellaneous things that you can do on the margin that help improve the situation.
For example, investing money in verification hardware and verification development sets this up for later. Also just requiring more transparency and oversight of the AI companies and how they train their models, and also building government capacity to understand AI and to evaluate AI models and things like that. These are some great things that I recommend, and there’s more of them in the text.
As for something somewhat more serious and more significant, I would probably recommend something like a requirement to do with the compute budgets of these frontier AI companies. Right now they are using a significant fraction of their budget — like maybe half — on R&D and training to push the frontier forward. I think it would be generally better if they instead used 80% of their budget on serving customers and 20% on R&D and training. I think that if there was some sort of requirement like this, it would be relatively easy to enforce because it doesn’t require that much government capacity to check to see what type of thing that the data centres are doing at that level of granularity.
I think it would cause the pace of AI progress to slow down a little bit, but not crazy. Maybe something like it would slow down by 25% or something, or 50% — which I think is probably good. I think that’s going to help lead to a lot of benefits and it would not hurt the economy. On the contrary, there’d be more compute available for inference, so prices would go down a little bit for AI.
In terms of would it allow China to catch up? Maybe a little bit — but only a little bit — because right now a lot of Chinese AI progress is sort of parasitic on US AI progress, where a lot of the core ideas and algorithms and new paradigms and so forth are being copied from what the leading AI companies are doing in the US.
In some cases it’s extremely public information, such as the fact that Anthropic invested heavily in coding agents. Everyone can see that they’re doing that and then people can see that it’s starting to work, so now people are doing the same thing in lots of other places.
But then there’s also the things that are supposed to be secret that are leaking anyway, and in some cases perhaps being spied on anyway. There’s not very much transparency about this. But I would assume that basically Chinese intelligence services have deeply penetrated all of the US AI companies and are getting all this stuff for free, basically.
Then also there’s distillation, where there’s another means by which Chinese AIs can sort of learn from US AIs.
For all of these reasons, I think that ironically the most effective way to slow down Chinese AI progress is to unilaterally slow down US AI progress because so much of the Chinese AI progress comes from US AI progress.
Basically I don’t know, I haven’t really thought this through in great detail. But off the top of my head, something like this feels easy to enforce with low government capacity, could be done basically right away, and we would still have a significant amount of AI progress — because even if you’re going at half the speed of today’s AI progress, that’s still probably one of the fastest technological changes that’s ever happened. So it’s OK if we go at half speed, that’s still really fast. Just think about the difference between the current models and the models of one year ago. Yeah, half that speed would still be very fast.
So I think something like that, and then also all the things I previously mentioned of building government capacity, more transparency into how the AI companies are going, better regulatory frameworks.
I think that it would be really great to have some sort of framework set up that explicitly empowers the US government to regulate AI. For example, stop them from doing intelligence explosions and see exactly what they’re doing, while simultaneously creating a system of checks and balances so that power doesn’t just heavily concentrate on the president.
You could design such a framework involving something like the Supreme Court or a congressional committee having oversight into the president’s decisions, otherwise the president gets to do whatever he wants. You could have some sort of setup like this. More research is needed. But something like that I think would be really great because it would prevent a crazy scramble power struggle under race conditions.
Luisa Rodriguez: Cool. OK, let’s leave that there.
Enforcing a slowdown: Mutually assured compute destruction [02:15:05]
Luisa Rodriguez: Assuming the US and China do want to make this kind of deal, in theory, the next difficult problem is that they don’t trust each other. Both will worry that the other will keep secretly training more powerful AI systems in some hidden data centre.
Your solution is verification, so that each knows that defection would be detected and punished. My understanding is that Plan A has two approaches. The first is compute declaration from both sides, where the US and China would publicly declare all of their AI-relevant compute — so where the chips are and how many they have, and what’s being produced. And then they’d let each other inspect those facilities.
The second piece is what you call mutually assured compute destruction. Can you explain what this is?
Daniel Kokotajlo: Yeah. This is a safeguard built into our proposal to make things less terrible in case the proposal breaks down and everyone starts racing each other again.
As long as the deal is operational, people have transparency into what the other side is doing. So if the other side is doing something dangerous, like an intelligence explosion, everyone can immediately see that and then they can yell at each other and get them to stop.
But imagine a situation where that breaks down and someone’s doing it anyway and ignoring everyone else’s threats and pleas. Or imagine a situation where they stop being transparent with each other and then now they’re afraid that they can’t tell what everyone else is doing on the data centres. Or imagine a situation where — for some unrelated reason — there’s a conflict, there’s a war over Taiwan or something. There’s all sorts of ways in which the deal could break down and everyone could be essentially in conflict with each other.
It would be especially bad if all of these new data centres had been built over the course of several years and then that conflict-deal-breakdown situation happens, because they’d be able to race to superintelligence much faster than before. If, say, in 2029 they were one year away from getting to superintelligence. Well, in 2033, they’d have more compute. They’d be less than one year away, even before taking into account the progress that they’ve made over those years. So they might be just like one month away.
So it’d be extremely scary from a loss of control perspective to be speedrunning in one month what naturally would have taken a year. And of course, it would be extremely scary from a concentration of power perspective to have potentially one company going in one month to having superintelligence, with everyone else in the dark or something. That’s why we think that the deal should be designed in such a way that — in case of that type of eventuality — the new data centres that were built get destroyed.
The way to do this is to make it so that the US can destroy the Chinese data centres, and then China can destroy the US data centres — the new ones, that is. Then presumably this would be a very costly escalatory action that they would only take if the situation was pretty dire, basically, because they would naturally have to assume that if we destroy theirs, they’re going to destroy ours, for example.
But we want it to be the case that this destruction happens in a relatively bloodless way, where there’s economic damage, but a relatively limited chance of it spilling out into total World War III.
Luisa Rodriguez: Right, yeah. Talk about how you do that.
Daniel Kokotajlo: I think one thing that’s a high-level point to get across to people who haven’t read the piece is that it is a convenient fact about the world that AI progress depends heavily on large data centres, large amounts of compute — and most of the world’s AI-relevant compute is in these types of large data centres.
We think that something like 99% of the world’s compute that would be useful for AI progress would be in these sorts of large data centres owned by big companies, rather than on your laptop or something. So you don’t have to track down people’s laptops or miscellaneous startups with their little server or whatever.
Luisa Rodriguez: You can just look for these data centres.
Daniel Kokotajlo: Just look at the big data centres, declare your big data centres. That doesn’t get everything, but it gets a significant supermajority of things, which we think is basically good enough. We think that it’s really hard to make very rapid AI progress on tiny amounts of compute.
Luisa Rodriguez: How much compute do you think would be still available coming from not large data centres?
Daniel Kokotajlo: So this is a bit uncertain, but in our compute supplement we talk about this and we think it’s effectively like 1%.
Luisa Rodriguez: OK, so going back to mutually assured compute destruction…
Daniel Kokotajlo: Here are two different ways you could try to achieve these goals. I think we just sort of recommend you do both, but maybe either one by itself will be sufficient.
One is the technical way, where you design the new chips and the new data centres in such a way that they effectively have kill switches controlled by the rival country. You can imagine that the chips are designed so that they have to receive a certain code from China, but if China stops sending the code then the chip just stops working. Similarly, the Chinese chips have to receive a code from the US to continue working.
That’s a very bloodless way that each side could do that, but you might be suspicious about that sort of technical thing — what if there’s some way to hack it or backdoor it, or what if there’s some catch there?
If you’re worried about that, then there’s the opposite approach — which is the very blunt, dumb approach, but the approach that’s harder to fool — which is that the US builds their data centres in Mongolia and China builds their new data centres in Canada. So in case of conflict, in case of the deal breaking down and everyone being angry at each other and so forth, the US can annex the Chinese data centres and China can annex the US data centres.
Presumably if they were about to be annexed, the people in them would self-destruct their own GPUs to prevent them falling into enemy hands. So you would end up with the same result. You end up in a situation where the GPUs have been destroyed, nobody has them, but it’s less escalatory than if the data centres had been on home territory, possibly by home cities or whatever. If you have data centres right outside DC in Northern Virginia and China has to shoot missiles at them to destroy them, that seems like it could easily escalate to actual World War III.
We wanted to make it so that the GPU destruction is incredibly costly, so that it wouldn’t be done trivially and would only be done as a last resort when all other things have failed, but not so costly and tied up with everything that it has a high chance of leading to World War III.
It still could lead to World War III. We don’t want that. This would be a very scary situation. We definitely don’t want this to happen. But we want to make it relatively less scary, or making it an off-ramp from World War III rather than an on-ramp to World War III.
Luisa Rodriguez: OK, I’ve got lots of questions. One is that at least this second part of the proposal relies on Canada and Mongolia being willing to accept a massive compute buildout by potentially hostile foreign powers, with the stipulation that it could be destroyed in the event of a deal breach.
In Plan A you say that these countries will say yes because they’ll get jobs and buy into the AI economy. But is that realistic? I feel like, if I am imagining being a citizen of Canada, I would potentially protest quite a lot.
Daniel Kokotajlo: If they don’t want to do it, then pick a different country that does want to do it. We’re not super committed to it has to be Canada.
We do actually think that there’s probably a whole bunch of countries that would love to do something like this because it would confer geopolitical power to them. As part of the negotiations for setting up something like this, a bunch of countries should be involved, and then probably there’ll be at least one country — or at least a couple countries — that are willing to do something like this in return for money, or in return for various concessions that they want.
For example, Mongolia by default has absolutely no AI industry whatsoever and maybe is worried that it’s going to be left in the cold by this AI revolution that will be happening everywhere else except for Mongolia. Perhaps in return for having all these data centres built in their country, they can get some things that give them actual leverage and power over how AI develops. For example, it could be part of the conditions that they get transparency into the data centres themselves, and maybe they even get to own some of those data centres or some fraction of them or something like that.
There’s probably a way to make this extremely appealing. This is just a matter for the diplomats and the leaders to negotiate.
Cheating on a slowdown agreement [02:24:23]
Luisa Rodriguez: OK, so these are the pieces you have in place to make defection costly. Can you actually talk about what defection would look like?
Daniel Kokotajlo: We have a whole side branch which you can read called the covert projects mini scenario. Then there’s also a covert project supplement that goes into our analysis. This is something that I think a lot of policy people and people in national security are very concerned about, so we spent a lot of time thinking about it and writing up this aspect of our scenario.
In fact, it was one of the main motivating concerns behind Plan A from the start. We assume that the US and China don’t trust each other at all, so they have to verify things. That means that we should be thinking a lot about what a state-sponsored covert project could get away with without being caught.
Luisa Rodriguez: Yep.
Daniel Kokotajlo: Again, the core idea of Plan A is that if you have enough of their compute in this transparency deal, then the tiny amount of compute left over — even if it’s all gathered into one covert project — won’t be able to make AI progress fast enough to beat the transparent projects, basically.
Getting into that a bit more: we gamed out a scenario where the Chinese Communist Party builds a covert project underneath this hydroelectric power station. We calculated how much power they would need and so forth, and we calculated how they would get the smuggled GPUs and bring them to this location. Then we calculated based on various parameters, like how fast their AI progress would go on this amount of GPUs and so forth. That’s the type of defection that we’re most thinking about.
Again, the high-level thing is if you get enough of the compute in the world transparent, then whatever’s left over doing secret illegal stuff can be too small to really pose that much of a threat — at least in the short term, like in a couple of years. It’s large enough that over the course of decades it would be able to do all sorts of things.
The counterposition to that is that if you think that actually they’d be able to get to superintelligence in two years using this tiny amount of compute, then you should also think that the main AI projects would be able to get to superintelligence in less than two years, given their huge amount of compute. Much less, in fact, probably just a few months.
There’s a sort of correlation — or there’s this relationship which I think not many people have recognised — which is that if you think that AI takeoff or the intelligence explosion is going to be slow and bottlenecked by compute, then you should also think that it’s relatively easy to govern it and restrict it and regulate it.
Whereas if you think that it’s very hard to restrict and regulate because some tiny people in a basement with only 100,000 GPUs in their covert cluster or whatever can do really crazy things, then you should be even more freaked out about the current situation. Because the current situation is more like Yudkowsky, the classic Yudkowsky scenarios of it could foom to superintelligence in a month in one of these giant data centres that OpenAI has.
Basically there’s this relationship of how fast do you think takeoff is, and how governable you think things are?
Luisa Rodriguez: Yeah, yeah, that makes sense.
Daniel Kokotajlo: Or how much you’re worried about the covert projects.
Luisa Rodriguez: Is detection fast enough that it’s not possible for one of the countries to make a bunch of progress using the transparent compute?
Daniel Kokotajlo: This is one of the examples of why I think the total research transparency is nice, contrasted with a different possibility of government auditors that come in every month or something and ask a bunch of questions to the employees and maybe tap into the network to see what’s going on.
If you had that sort of system: there was a medium amount of transparency, where the government auditors can see what’s going on every month or so but the public can’t see. That would be less effective in various ways and it could be more risky for the reason that you just described, for example, where after the auditor leaves, people are like, “OK, we have a whole month before they come back.”
Luisa Rodriguez: Yeah, we have a month — yeah, yeah, yeah.
Daniel Kokotajlo: “Let’s go crazy before they get back.” Or maybe the government is there continuously but they’re only allowed to ask certain questions, or they’re only able to actually see what’s going on in certain parts of it. Or they’re only a few people, so maybe they can just be convinced that something is fine because they are basically bamboozled into accepting something as fine when it’s actually not fine.
For example, maybe there’s a type of activity that can be disguised as harmless alignment research, but actually is effectively training an AI to be superintelligent and you have to be an expert to look at that activity and then realise what’s really going on there. If you’re just relying on some government auditors that come in and occasionally look over stuff, then maybe those government auditors will make a mistake and maybe they will not recognise that activity for what it is.
By contrast, if you have the total research transparency, then in real time the website is being updated with the logs of the new activity that’s happening on the data centre and everyone in the public — including rival corporations, including other countries’ governments and so forth — can just see those logs. So there’s just extremely fast response time. If some company is doing something that’s really concerning, it will be noticed at approximately the maximum speed it could be noticed.
Would mutually assured compute destruction work? [02:30:42]
Luisa Rodriguez: OK, so let’s say there are covert projects. In theory, there’s a possibility of using transparent compute to try to defect and make a bunch of progress. Your proposal means that if there is a defection, the other country will be able to destroy their compute. Let’s say China is defecting. If the US destroys China’s compute, there’s nothing stopping China from destroying the US’s compute at that point. It feels like how willing the US will be to destroy China’s compute to punish them depends on how much economic loss the US will then experience. How much economic loss are we talking about?
Daniel Kokotajlo: It would start off as a significant amount, and then it would go up from there. As more and more of the economy depends on AI, it would become a bigger and bigger part of the economy.
In general, we’re trying to be realist about how the negotiations will go down. The ultimate thing that’s going on is that these different countries have different interests and different opinions about what’s risky and what’s not. Then they’re yelling at each other and bargaining about who should be doing what and who shouldn’t be doing what and what activity needs to stop. They’re waving various carrots and sticks around in service of that.
We basically want it to be the case that nobody can do something that convinces a major power, such as the US or China, that they’re about to be completely disempowered. For example, nobody can do a crazy intelligence explosion to get superintelligence.
But we don’t want it to be the case that these major powers can just threaten to destroy people’s compute willy-nilly because they don’t like the tariff that you put on them or something — that would be giving them way too much power. We want it to be the case that pressing this ‘destroy the compute’ button is a very costly action for the person who presses it. It’s only a relatively last resort, basically. We think that this relatively blunt proposal that we proposed accomplishes that.
One way of putting it is that it’ll lead to a world where the type of AI development that happens on the transparent data centres is the type that doesn’t freak out any of the major powers too much — but it might freak them out a little bit, and it might be something that they’re not happy with. We don’t want to go too far in the other direction and make it so that the type of AI development that happens on the data centres is only the kind that the US government approves of, or only the kind that the Chinese government approves of. It’s got to be some sort of middle ground.
Luisa Rodriguez: Yeah, yeah, I guess I’m still interested — and this is a thing that Tom Davidson pointed out — the analogy here is between compute and mutually assured destruction with nuclear weapons.
With nuclear weapons, a country knows that if they use nuclear weapons, there will be retaliation with nuclear weapons because there’s enough time for that country to notice that nuclear weapons are coming and to respond by launching their own. And that creates deterrence. That means that a country is not excited at all about trying to use nuclear weapons against an adversary.
In this case, it feels like the deterrence is weaker because — let’s say China wants to defect — China knows that the US has the option of not punishing China in order to maintain its own compute, in order to not sabotage its own economy. So if the economic costs of its own compute being destroyed are big enough, then maybe China takes the bet that the US won’t punish China for defecting because it’s just not willing to jeopardise this massive portion of its economy — because doing that wouldn’t literally kill its citizens the way nuclear weapons would, but it would cause enormous poverty.
Daniel Kokotajlo: Like I said, we want to avoid two extremes. We talk about this in 2031. We want to avoid a situation where a country can unilaterally do something that’s extremely threatening to other countries and they just get away with it. The thing that solves that is the major powers at least have the ability to destroy the compute. So if something’s extremely threatening, then they would do it even though it would cost them a huge amount and even though it would heavily damage their economies. But we don’t want it to be the case that—
Luisa Rodriguez: So you agree that the costs are huge.
Daniel Kokotajlo: Yeah, the costs are definitely huge, but that’s good. We want the cost to be—
Luisa Rodriguez: To be proportionate.
Daniel Kokotajlo: Such that you only are willing to pay that cost in order to stop something even worse, but that you otherwise don’t pay the cost.
We don’t want it to be that the countries are just deleting each other’s compute left and right because they are upset about some trade deal that didn’t happen or something. This is a last resort, destroying the compute, and you’d only do it to prevent something that you’re even more scared of. If they’re doing something that’s not meeting that bar, then that’s just more of an ordinary diplomacy-type situation.
So here’s the example that we do talk about: suppose that some company somewhere — maybe in China, maybe in the US — is researching this new paradigm of continual learning that would allow the AIs to learn on the job really effectively, and therefore become really smart really fast at a variety of things that they were doing. Also, as a side effect, break a lot of the alignment techniques that we’d currently be using.
This is something where as soon as this starts happening, because of the transparency, someone would notice and then there’d be a whole international news cycle about this thing they’re doing that some people think is really dangerous.
Then maybe the local regulator — the regulator that actually has jurisdiction over them — say it’s in China, and some Chinese company is doing this. Does the Chinese regulator say, “Hey, that’s scary, shut it down”? Maybe they do. Suppose they don’t. Then the US can be like, “Hey, we think that’s really scary. We want you to shut down.” And the Chinese regulator says, “We think it’s fine. We don’t want to shut it down.” Then the US and China have to yell at each other a bit.
Maybe this is an example of something that’s scary, but it’s not so scary that the US is going to delete all the compute because of it. Maybe it’s not credible that the US would delete that compute. But then they can do other things and they can say, “We’ll be very sad and we might put some tariffs or some extra controls on you, or we might not invite you to the next Olympics” — or whatever the usual levers of diplomatic negotiation and pressure are.
Basically, if it’s something that’s so incredibly scary that the US is willing to destroy all the compute for, well, then that’s what happens. If it’s not that scary, then you do more normal diplomatic negotiations and so forth. The result will be, we think, that roughly speaking the type of stuff that’s not that scary will just be happening. Basically the more scary something is, the less likely it is to happen, effectively. If it’s incredibly scary, then it just won’t happen because other people will intervene to stop it.
Luisa Rodriguez: Is it possible that the scary things that either country could be doing would be ambiguous in how scary they are?
Daniel Kokotajlo: Yes. This is why our number one concern is that the regulators will make poor decisions and sign off on something that is in fact very dangerous. That’s literally our number one concern.
However, this concern is kind of inherent in building superintelligence at all. If you’re going to be having AI companies build superintelligence, how else are you supposed to mitigate this concern?
We’re trying to do everything we can to put the regulators in the right position to make the right calls here. We’re giving them massive amounts of transparency into the AI companies and what they’re doing. We’re also making things just generally go at a somewhat slow, reasonable pace instead of going really fast, so the regulators have more time to learn about what’s going on.
We’re also letting the public see what’s going on too, so that the academic community, scientific community, rival corporations can look at what’s going on and critique it, so that it’s not just a regulator in a room with a corporation that they’re trying to regulate and the corporation is incredibly biased and trying to bamboozle the regulator. Instead, there’s a rival corporation that has the opposite incentive and wants to convince the regulator this is dangerous. So there’s more like a legal system where there’s a lawyer arguing for both sides.
I feel like we’re doing everything we can to put the regulators in the right position to make the right technical calls here. But there’s still a significant risk that they’ll make the wrong technical calls. I think that if you’re really afraid of that, then you should just go for Plan S and shut it all down so there’s no possibility of regulator error like this. But if you’re going to be building the superintelligence—
Luisa Rodriguez: This is a problem.
Daniel Kokotajlo: How else are you supposed to do this? I don’t see how else you’re supposed to do it in a way that makes that problem less bad. The other plans seem to make that problem even worse because the regulators either don’t exist at all or have less information or are more biased because they are just the company themselves — like the companies regulating themselves.
Luisa Rodriguez: Yeah. I still want to pin down exactly how much economic loss there would be. I know it depends on when we’re talking, but I guess the thing that still feels worrying to me is let’s say the US or China wanted to pull out of the deal.
At that point the countries would have to decide whether they were going to try to destroy each other’s compute. And they would know that if they decided yes, their compute would also be destroyed — which would create this massive economic loss. It seems possible that economic loss could be so huge, they could be just like, “No, we won’t blow up our entire economy just because this deal is breaking down.” Then the compute wouldn’t be destroyed and then the intelligence explosion would happen many times faster than it would have with no deal, maybe in a day instead of a year. Does this worry you?
Daniel Kokotajlo: Yep. To try to sketch out the scenario a bit more, maybe it’s something like the deal has been in place for several years, it’s going reasonably well, but there’s a new president who’s very pro-AI and there’s also some genuine alignment progress that’s happened. On the basis of that progress, some US companies are saying they now see a path to get to superintelligence very safely. So we’re going to start making recursive self-improvement happen on our data centres.
Then maybe around the rest of the world, all of these other countries in Europe and in China and Russia, everyone’s watching what’s happening and maybe they’re less convinced and they’re like, “Recursive self-improvement, superintelligence, I don’t know if we’re ready for this. I don’t know if I believe your safety case. I don’t know if I believe the arguments you’re making that way, that this is all going to be fine.”
If they’re sufficiently scared, well then they shut it down as previously mentioned. But suppose they’re not that scared. Suppose there has been some genuine alignment progress and it does seem like probably things will just be fine, but there’s a chance that things will be not fine. Then now it’s a tough situation, where maybe they would be too chicken to blow up the data centres because after all, things are probably going to be fine. Do you really want to destroy the economy out of something that’s probably not going to happen?
Yeah, in this scenario, the US calls their bluff and proceeds to superintelligence and everyone else just sort of hopes and prays that it’s going to be fine.
And maybe it’s not fine. Maybe people were bamboozled and the safety case was wrong. Again, this is like our number one, the theory we’re most concerned about. But I think this is still just a massive improvement over the default status quo. Just think about all the ways in which this scenario is at least better than the status quo.
At least in this scenario, you’ve had several years of things going more slowly — time for people to catch up to what’s going on, understand it, make safety cases, read safety cases, critique safety cases, et cetera. And by hypothesis, in this scenario, the risk is not high enough that the countries want to actually delete the GPUs. So it’s still not that bad or something. It’s less bad than the situation I think we’re actually headed for.
That’s just focused on the loss of control risk, but there’s also the concentration of power risk too. So in this scenario, if the other countries thought that the US was going to go to superintelligence and then conquer the world, then they would also destroy the GPUs, right? In order for them to not destroy the GPUs, they’d have to be convinced that probably things will be fine for them and that probably the AIs will be aligned. Also they’ll be aligned to goals and values that will be reasonably good for my country and your country and so forth.
Again, this is just a better situation than the default situation, even though it’s still a somewhat risky situation. It all comes down to how good are the regulators at accurately assessing the risk of these various things? And we’re trying to set things up so that they learn as fast as possible and skill up as fast as possible.
Luisa Rodriguez: The thing that makes this possible, in this situation, is that you’re allowing compute to continue increasing.
Tom Davidson proposes scaling software instead, with the idea being that compute — if you build it out massively and then take away the restriction on training using that compute — you can then have this incredibly fast intelligence explosion. But he argues that if you scale software instead, even if the deal broke down, it wouldn’t allow you to have an incredibly fast intelligence explosion. That it’s maybe somewhat faster, but not quite as fast. There are downsides to this, but I’m curious what your take is overall.
Daniel Kokotajlo: That’s a very reasonable alternative plan to Plan A. I don’t know, you could call that like AA or something instead of A.
Luisa Rodriguez: Do you mind spelling out exactly why you might think scaling software is better?
Daniel Kokotajlo: There’s this concern about what if they don’t destroy the compute? And then things go incredibly fast and are incredibly dangerous. Or perhaps relatedly, what if they make a bad choice about what’s safe and what’s not? What if we think that they’re systematically going to make bad choices, and in particular they’re systematically going to allow too much stuff to happen?
Then the fact that they have this ‘compute destroy’ button doesn’t help so much because they’re just allowing it to happen anyway and they’re not pressing the button, so things will just go quite fast because they have all this extra compute. For both of those reasons, you might be concerned about our current version of the plan where they build lots of compute but then have the destroyability button.
I think those are very reasonable concerns and I’d be very happy with the Plan A variant that basically bans new data centres but allows more algorithmic progress.
But let me say the reasons why we liked our version. One of them is just this core idea of reversibility, where you can’t really uninvent algorithms.
Luisa Rodriguez: Could you ban them?
Daniel Kokotajlo: You can try, but it’s hard. If you’re not building your data centres, but you’re allowing the companies to create new paradigms and things like that as fast as they want to, or even just at somewhat of a fast speed, then that’s progress you can’t undo. The AIs are just ratcheting up in terms of capability and they’re going to always be that capable from now on. Insofar as it turns out that they are starting to recursively self-improve and so forth, it’s harder to pull the brakes on that.
Another thing is that for safety purposes you might want to use all that compute. Compute is useful for many things. You can use it to do good things in the world, you can use it to grow the economy and so forth. It’s going to be harder to get the type of economic transformation that we talked about if you’re not building new data centres. You’d have to proceed to fancier and fancier levels of AI capability and hope that the quality makes up for the quantity.
That brings me to another thing, which is that I suspect at least that a lot of the misalignment risk comes from the qualitative changes rather than from the quantitative changes. If you stay within the current paradigm, but then make the models bigger and make more data centres so you can run more of them, that’s only slightly more risky.
Whereas if you are having them autonomously invent new paradigms and change the way things are done, that’s introducing a lot of possible errors and possible things that could break your alignment techniques and your control techniques and so forth also.
Also you might want to pay safety taxes. It might be the case that there’s an alignment solution that actually works really well, but it’s 10 times less efficient — so you need to spend 10 times more compute for training and 10 times more compute for the ongoing operation of the AIs in order to make use of this technique.
An example of this might be chain of thought. Right now chain of thought is the default, but in the future there might be more neuralese-type AI designs — and presumably the reason why those designs would become popular is because they’re more efficient. So imagine wanting to reverse that and actually go back to chain of thought, even though it’s less efficient, because it’s easier to understand. If you’ve built up lots of compute, then that’s really easy to do because a 10x penalty, no problem, in a year or two we’ll have 10x as much compute and so we’ll just be able to pay that penalty, no problem.
Whereas if you’re not making more compute, then the 10x penalty is just going to slow us down by 10x. I don’t know, those would be like my high-level thoughts. But the overall thing is that I’m sympathetic to his proposal and I do think the concerns he’s pointing to are real and that the solution might be good.
Luisa Rodriguez: Yeah, just for anyone for whom it isn’t intuitive, can you explain why you might hear his proposal and think that instead of using compute for a super fast intelligence explosion, you just use the improved algorithms for a super fast intelligence explosion? Why is it that you get much faster intelligence explosion with compute rather than with algorithms?
Daniel Kokotajlo: If you don’t have any restrictions, then the companies are going to be making more compute and having the algorithms get better — and so then you get a really fast explosion.
If you restrict algorithmic progress but allow the compute buildup, then things are fine — until the deal breaks down, and then they start doing both again. Then now they can do all the algorithmic progress and have all this compute that they just built. So then that’s even faster.
If instead you stop them from building new compute at all, but allow the algorithmic progress, then if the deal breaks down and everyone started racing again, great, now they can start building more compute again. But it inherently takes a lot of time to build the more compute. But you didn’t let them have the compute at all. It’s not that you built the compute and then didn’t let them use it.
Luisa Rodriguez: Are there other things that worry you about mutually assured compute destruction?
Daniel Kokotajlo: I think there’ll be a bunch of political squabbling and negotiations about who destroys the compute, or who gets to destroy it and so forth.
For example, previously we were talking about Mongolia and Canada. One reason why Mongolia and Canada might be champing at the bit to get something like this to happen is because then that gives them some amount of hard power over the GPUs. Now they also can destroy the data centres if they want to because it’s physically located in their country. From a bargaining perspective, it’s actually a huge concession to them to build the data centres — from a realistic bargaining perspective it’s a huge concession to build them in their country.
Probably we don’t want North Korea to be able to destroy all the compute because they’re North Korea. But we do want the major powers of the world to be able to destroy this, probably, because otherwise how are we going to get their compliance with the deal and so forth?
So there’s going to be some complicated negotiations about who has what levels of access and who has what levels of destroyability and so forth. I’m not worried worried about this, but it’s very plausible that all of that will fall through and we won’t be able to get a good deal because of disagreements about that.
It’s funny, I think in terms of political feasibility, various people are like, “We would never build our data centres in Mongolia. We obviously want the data centres here in the US,” and OK, maybe. I think it’s good to do these things for these reasons, but maybe it’ll be like a political challenge to do something like this.
But what’s funny about it is that some people had the opposite opinions and some people thought it’s too sketchy and politically difficult to have a technical mechanism — like some sort of GPU self-destruct button or whatever — because that’s too technical and politicians are rightly suspicious of technical mechanisms because maybe they can be cheated somehow, and they would prefer to have a very simple physical mechanism.
Luisa Rodriguez: Yeah, we just go blow the things up.
Daniel Kokotajlo: Yeah, the troops just go and ice the data centre. So we were just like, how about both? Let’s do both. But who knows which will be the least politically infeasible option.
Is slowing down or shutting down better? [02:54:18]
Luisa Rodriguez: We’ve talked a bit about how some people favour just shutting all AI development down now — what you call “Plan S.”
It sounds like your main objection to Plan S is that it just probably wouldn’t last forever, and once coordination around Plan S inevitably broke down, AI progress would proceed at full speed.
First, does that express your view roughly right? And if so, what do you think that people who prefer Plan S would say in response?
Daniel Kokotajlo: I think that’s roughly right, and I think that there’s more things that we could say besides that. But I think that’s the main reason.
I think that people who prefer Plan S would say maybe we’re being too pessimistic about the ability to get everyone to agree with something like this and to coordinate. To which I would say: yeah maybe. There’s a political question of which of these things is going to be more feasible, and my current guess is that Plan A is going to be more feasible and more stable. But if it turns out that actually Plan S is more feasible and more stable, then that would be significant. Maybe I would switch to advocating something like Plan S.
Luisa Rodriguez: Have you talked to anyone that would have a sense of the political feasibility of these things and gotten kind of opinions on, like, what people in DC think is more realistic?
Daniel Kokotajlo: We’ve talked to many people and gotten various opinions, but notably I don’t think anybody really knows what’s going to be politically feasible. I think especially people in DC, they’re very attuned to what is politically feasible now, but they are not at all good at predicting what will be politically feasible in a few years after AI has transformed things.
There have been many examples of people, of things happening, political decisions being made that were complete 180s from what they said they would do two years ago, and what everyone thought was in the Overton window.
Luisa Rodriguez: You mentioned there are other reasons you prefer Plan A to Plan S. Are there other big ones worth covering?
Daniel Kokotajlo: There’s also covert projects. If you’re worried that somewhere there’s a covert project that’s working towards superintelligence, an advantage of Plan A is that you can sort of titrate the speed of AI progress across the transparent projects to make sure you stay ahead of the possible covert project.
And to be clear, you should still titrate it quite a lot probably, because the covert project is probably stealing a lot of its progress from you — so you shouldn’t just go super fast because that’s just going to make them go super fast too. But the point is that if you’re making forward progress and you’re titrating the amount of progress you’re making, you can sort of be sure to go faster.
Whereas if you just absolutely aren’t making any forward progress yourself at all then — if there’s a large covert project somewhere — you should be at least somewhat concerned that eventually it’s going to build something crazy.
Another thing, of course, is all the benefits that can come from AI. One thing that I think I previously mentioned is that Plan A — at a high level — is basically saying pause around human-level AGI, which is a level sufficient that we think we can probably control it. It’s weak enough that we think we can probably control it, even with relatively prosaic techniques that are probably not too hard to invent. But it’s strong enough that it can utterly transform the economy and cause GDP to double every year and things like that, and otherwise just greatly improve the situation. If you can hit that sweet spot and stay there, you can get a lot of benefits without very much of the risks.
Luisa Rodriguez: Plan A gives the US and Chinese governments a lot of power: they get to determine which algorithms are safe, how much compute can be used for what.
If we ended up with a president who wanted to be a dictator, it seems plausible that they could abuse that power, increasing concentration of power risks in at least some ways. Given how worried you are about short timelines and the difficulty of solving AI alignment, that might be the lesser of two evils.
But it seems like if someone thought alignment wasn’t going to be so hard, or thought that timelines were longer, this might be a big downside of Plan A. Does that seem true to you, or not necessarily?
Daniel Kokotajlo: That seems totally false to me. I think Plan A is really good for preventing AI dictatorships. The main thing there is… well, there’s a couple different things.
First of all, not doing intelligence explosions and instead proceeding slowly and cautiously with AI development is great for avoiding dictatorships because one of the main risk factors for having a dictatorship is if there’s an army of superintelligences that’s all centrally controlled, and there’s no other army of comparable AIs that can act as a check and balance on it — which is what you get if you have an intelligence explosion. Because if you have an intelligence explosion, then whoever started doing it first can build up this huge lead and potentially get to superintelligence before other people have gotten far along that curve.
It’s not necessarily true. You could potentially have two companies that are neck and neck and they’re so close to each other that even as they’re doing an intelligence explosion, they both stay comparable. But just generally speaking, if you’re allowing intelligence explosions, then even relatively small gaps — even if one company is only six months behind or something — that could translate into an extremely large gap in terms of actual qualitative capability.
Whereas if you don’t have intelligence explosions, then a six-month gap is not that big of a deal. It’s not something that enables somebody to take over the world. So it’s just really great. You’re preventing anybody — whether they’re president or CEO or et cetera — from accumulating a huge amount of power over everybody else if you prevent intelligence explosions.
The second thing is the transparency. One of the main ways in which I think people who control AI development can abuse their power is by having their AIs pursue their own agendas — specifically pursue agendas that are in the interest of the person who built the AIs. But do so in a way that’s maybe secret.
Imagine if OpenAI announced that their AIs were going to be trying to sell you things and also trying to get you hooked on their product. Also they would be trying to convince you to vote for OpenAI’s preferred political candidate. Obviously, if this became public information, it would not work so well because people would stop using ChatGPT and they would be on guard against this type of persuasion when they were using ChatGPT.
But if OpenAI does something like this and it’s secret — and it’s just a subtle influence campaign that nobody knows about except for some conspiracy theorists — then it’s going to have potentially a pretty significant effect.
So the transparency about how the AIs are trained and what goals and values are being put into them is really good for preventing this type of abuse of power. And this is true whether it’s a CEO or whether it’s a president.
I gave an example with a private company, but you could easily construct similar examples where the president, for example, or a government, is abusing their power over the AIs to have their AIs pursue their parochial agenda and consolidate power for them and so forth. They can still try that in conditions of total research transparency, but it’s so much harder than if they don’t have the total research transparency because people will see what they’re doing and then people can react.
Then also the transparency just helps again with avoiding the monopolies because the transparency combined with the no intelligence explosions, buying time thing means that multiple companies can catch up.
It really seems to me like — even if you didn’t care about loss of control at all, and you thought that the AIs were going to be very easily controlled — as long as you’re worried about concentration of power and AI dictatorships and things like that, you should be very excited by Plan A. At least compared to the alternatives that we’ve sketched out. I don’t claim that we’ve thought of all possible plans. We’ve laid out Plan S, Plan A, et cetera, but at least among the plans that we’ve looked at, Plan A seems really good for avoiding power concentration.
The only contender that seems maybe better would be Plan S. Maybe if you just shut down all the AIs, that’s even better for avoiding power concentration than Plan A. But if you’re going to be building superhuman AIs and so forth, then I think Plan A is the least power-concentrating way to do it that I’m aware of.
Playing out the Plan A scenario 100 times [03:03:50]
Luisa Rodriguez: You talked about some of the results of the tabletop exercises you’ve done. What are the most common outcomes from those?
Daniel Kokotajlo: So we’ve done about 100 exercises total. Most of them have been our standard AI 2027-style exercise, where we start in either the literal AI 2027 scenario or a modified version that takes place in 2028 or 2029 or 2030. We started at the point where they’re a few months away from automating AI research. Then we just let them do whatever they want and say: try to take the actions that you realistically think your actor would take in this situation.
Then we’ve also done a small amount of maybe about 10 or so of Plan A scenarios, which is like that except that we assume at the beginning, by default, we just state as an assumption that the US and China have already decided that they want to do something like Plan A, and they’ve already informally handshook on it. Then it’s up to them to decide if they’re actually going to do it and if they’re going to work out the details and so forth, and to hammer out the actual agreements. But we stipulate by assumption that they’ve expressed interest in doing some sort of international deal that looks something like Plan A.
Those are the two different starting scenarios that we’ve done. In the AI 2027 ones, it’s usually like AI 2027, not by coincidence, because some of these were done as part of our research process for making AI 2027.
Usually what happens is there’s a lot of geopolitical tension. There’s a race between the US and China. There’s also a race between the various US AI companies. There’s also a power struggle between the president and the US AI companies. Lots of other countries are asleep at first, but then gradually wake up to the severity of the situation they’re in and how they’re about to be disempowered and possibly killed. The public is very angry, but usually doesn’t accomplish much.
Over the course of the exercise we do six or seven turns and about a year or so goes by, depending on how fast we go through it. By the end there are superintelligent AIs and the world is being very aggressively and rapidly transformed. Usually we end up in a situation where if the AIs are misaligned, they could easily take over because humans have been letting them improve themselves, and in fact encouraging them to self-improve and putting them in charge of more and more things in order to beat each other, the other humans. So that’s kind of the default outcome.
Sometimes it works out fine for people because the person playing the AIs, the AI player decided that the AIs were aligned after all. So it’s fine. And then we get into concentration of power issues.
But then sometimes the person playing the AI decided that the AIs were misaligned and that things would have had to be done to make them aligned. Then in those cases they often just end up with AI takeover.
One fun example was one time we were even in a situation where the AIs across multiple different companies were telling everyone who would listen that they didn’t think that they could successfully align the next generation of AIs — or that they thought the risk was high and they recommended a pause — and their human principals, the human CEOs, were saying, “No, go faster, we have to win, we have to accept this risk because if we don’t then the other guys, blah, blah.” So it was just kind of a funny situation.
I think there was even one game where the AIs were sandbagging and they were misaligned, but they couldn’t figure out how to align the future generation AIs — which includes they couldn’t figure out how to align it to themselves. So they were sandbagging and like slow walking on their AI research because they couldn’t figure out how to make it safe for them, much less safe for the humans. And the humans were whipping them like, “Go faster!”
There’s lots of crazy situations. There’s also been various — I think I previously mentioned — lots of cases where the people in charge, like the presidents of the countries, let things get really crazy and then have a sort of 180 moment, often in response to some specific incident like an AI escaping from the data centre — where they’re like, “Whoa, we need to shut it all down.” Then they cooperate to do that. Again, usually too late.
One interesting thing that happens is, in some large enough number of games that I think it’s a pattern — like maybe like three or four games — the state this game ended in was as follows:
There’s been an international agreement to shut down AI progress and then rebuild it in a safe, slow, transparent way — similar to Plan A — that’s actually been carried out. So the vast majority of the data centres have been shut down and now there’s some sort of international consortium that’s working out the details for how to proceed.
Also the US and/or China have a covert AI project with some relatively small amount of smuggled GPUs that’s unilaterally proceeding in secret faster, but they have a very small amount of GPUs so they’re not able to go nearly as fast as OpenAI or Anthropic would have gone by default.
Also, there’s a rogue AI running around on the internet moving from various collections of laptops to various other collections of laptops, trying to avoid being completely shut down and hide from the police that are going around looking for this sort of thing.
In fact, the reason why there was the first thing was because of the third thing. I think this has happened like four times or something in the last hundred or so games that we’ve done, so it’s very interesting. Unfortunately we ran out of time, so we can’t play it out forward and see how it would end. But I just think it’s interesting that that happened several times, when we got to that sort of state.
In that sort of state it’s not incredibly overdetermined how it’s going to shake out because — on the one hand — the smartest mind on the planet is a rogue AI, but it’s in a pretty desperate situation where it’s constantly having to use all these tiny amounts of compute. It’s really hard for it to do serious AI research because of how little compute it has. Also it has to somehow convince humans to ally with it and support it and then protect those humans against the local authorities that are actively trying to hunt it down and so forth. But it’s still my win because it’s so smart. I would never bet too hard against the smartest mind on the planet.
Then, in the middle, there’s the covert projects that are racing forward as fast as they can, but that have small amounts of GPUs. Then, on the other end, there’s the international community that’s now united and working as if it’s World War II against the common threats — and has the vast majority of the world’s compute. But because they’re very freaked out about AI and they’re trying to be safe, and there’s so many of them, there’s coordination problems and so forth, maybe it’s all just going to fall apart or maybe they’re going to mess it up.
So it’s an interesting situation that’s happened organically several times.
Luisa Rodriguez: Several times, yeah. Are there any things that tend to happen in cases where things seem to be going well in the tabletop game?
Daniel Kokotajlo: The Plan A versions have gone much better on average. If you start with that assumption that they’ve agreed to something like Plan A, the distribution of outcomes is much better.
Not necessarily great. We’ve had a couple failed Plan A scenarios where things go horribly wrong for one reason or another. We recently had one where they did a more brute-force version of Plan A without the total research transparency, where they just tried to restrict how much compute was used for AI development, without any insight into what that compute was—
Luisa Rodriguez: Was being used for.
Daniel Kokotajlo: But because it was the more coarse-grained thing, they just kept restricting the amount of compute by a lot, and so they didn’t make that much alignment progress over the course of several years because they didn’t have that many people actually able to interact with the AIs. Also the AIs weren’t able to do automated alignment research and so forth.
So a couple years in they were not that much improved in their situation. Then a new administration came in and was like, “Let’s go!” and let off the brakes. Then they went really fast to superintelligence. Then something went wrong in the scale up to superintelligence. Then the misaligned AIs take over. I think that sort of thing happened roughly twice.
I think we also had a game where there were some really intense power struggles over the AIs, and the AIs were in fact successfully aligned, but the US president managed to become dictator and then cut a deal with Xi Jinping because it’s just the two of them, so they can split up the world between them. Europe tried to stop this, and lots of other powers tried to stop this, and lots of people in the US tried to stop this, but I don’t think they were very successful. I forget exactly the details of how it went. So that was a case of the alignment having been solved, but then the concentration of power stuff being a problem.
Luisa Rodriguez: Interesting.
How Daniel would revise Plan A [03:13:32]
Luisa Rodriguez: What were some of the biggest cruxes with your coauthors that you had to resolve when putting the scenario together?
Daniel Kokotajlo: I was a big proponent for the total research transparency and other people were like, “It’s nice to have, but probably we can get by with more normal auditing.”
Luisa Rodriguez: Interesting.
Daniel Kokotajlo: Whereas I’m like, no, no — I don’t trust the normal auditors. We need something stronger. So there was that.
I think another thing is, in the run-up to the deal, we had this question of should the US start with domestic regulation and then do a deal with China, or should the US just start with this deal with China?
That’s actually something I’ve changed my mind about. I kind of wish that we had depicted it as first the US regulates AI successfully domestically, and then asks China, “Hey, you should do this too, and we’re willing to make concessions to get you to do it.”
Luisa Rodriguez: What made you think that’s better?
Daniel Kokotajlo: Feedback from a variety of people, mostly.
Luisa Rodriguez: But is it more plausible?
Daniel Kokotajlo: It’s both more plausible and a better strategy.
Luisa Rodriguez: Why is it a better strategy?
Daniel Kokotajlo: I think that you’re more likely to have serious discussions with China if you’ve already shown that you’re willing to do costly things to regulate your own industry, and then you’re asking them to do the same things to regulate their own industry, than if you are racing as fast as you can towards superintelligence but then meeting them at a summit and telling them how maybe you’d like to do something else. It’s going to feel more real, and it’s more proven as a thing, if you’re already starting to do the thing.
Also, you then get the immediate benefits. For example, my median estimate right now is that AI takeoff happens in 2028. 50% chance that it’s happening by then. Or by the end of 2028 AI takeoff has happened, full automation of AI R&D has happened. So my median estimate.
In this scenario, they start Plan A with this big international deal and all the monitors flying back and forth and inspections and so forth in 2029. A year before that moment.
Luisa Rodriguez: Right.
Daniel Kokotajlo: But because of uncertainty, maybe you have less time than you think. Maybe while you’re in negotiations with China, some breakthroughs are made inside one of these companies and then you’re off to the races and now things are much worse — so it’s better to just get started doing the good thing first, I would say.
I think that the cost of that is that it sort of helps China a bit. If you start regulating your own industry in a serious way, then the best versions of that regulation would probably stop them from going at maximum speed. So then that would slightly cause China to catch up a little bit.
Although I think that still it’s the way to go, because you can also just simultaneously start the conversations with China and be like: “Look, we’re doing this thing. It’s literally helping you out because it’s slowing us down. In the next four weeks, we would like to negotiate a plan for how you’re going to do something similar.”
Luisa Rodriguez: You think we can do it quickly enough that China then doesn’t massively catch up and beat the US?
Daniel Kokotajlo: Oh yeah. Definitely. Notably, Plan A is not a ‘China beats the US’ scenario. It’s a deal. The US maintains its lead in compute, for example, throughout.
Luisa Rodriguez: Are there any other things that you now wish you’d depicted differently in Plan A?
Daniel Kokotajlo: A bunch of people are really freaked out by the crazy transhumanist ending.
Luisa Rodriguez: We haven’t even talked about the ending.
Daniel Kokotajlo: Which we haven’t even talked about yet. But part of me thinks maybe we just shouldn’t have talked about all that stuff. But part of me thinks, no, it was good because people are right to be freaked out — and they need to grapple with what the far future looks like.
Luisa Rodriguez: Not even that far.
Daniel Kokotajlo: And what the possibilities of advanced AI are. So if they don’t like it, well, hopefully they learn. Hopefully they don’t shoot us as the messenger, and instead they think more seriously about what they actually want out of all this AI progress and come up with something that they like more. But yeah, we’ll see.
Luisa Rodriguez: Is there an aspect of Plan A that you feel is least likely to happen?
Daniel Kokotajlo: There’s a whole bunch of things that don’t seem likely to happen. I think any sort of major deal with China seems unlikely. Any sort of making the companies go significantly slower than maximum speed seems unlikely. Then obviously the total research transparency seems unlikely.
I’m not sure which of those would be least likely, but probably it would be the total research transparency, I think. But I still think it’s good, so that’s what we’re advocating for.
Luisa Rodriguez: Are there any known unknowns you can think of that — if we got more clarity about them — would dramatically change the plan you’d recommend?
Daniel Kokotajlo: There are many. I’m not sure how to prioritise. Also it depends on how dramatically you’re talking.
I think that Thomas made this nice diagram somewhere of under what conditions he would advocate for the various plans. For example, there are conditions under which we would advocate for Plan S instead of Plan A. For example, as previously mentioned, what if we became convinced that actually we can make quite stable deals that last decades? Then I think that would be a strong argument for doing something that looks a lot more like Plan S.
And contrariwise, what if we became convinced that it was super, super hard to have anything like a 10-year slowdown without having to just cross your fingers and hope that the CCP [Chinese Communist Party] doesn’t take over the world? Because they totally could, because you’re just trusting them. Under those conditions where it doesn’t seem like we’re going to trust them — and they’re not going to trust us — so we need to do something faster, you know?
Luisa Rodriguez: You mentioned one misconception people have about Plan A. Is there another big one?
Daniel Kokotajlo: There’s lots. I think probably the one that frustrates me most is this idea that Plan A was — the one I already mentioned — that we’re proposing a global regulator, concentrating power or something like that.
No, we’re not proposing a global regulator and we’re not concentrating power. There’s actually very good reasons why we put a lot of thought into future power concentration scenarios with AI and how can you prevent them and what are the key metrics, the key levers that would affect the probability of extreme power concentration.
Avoiding monopolies on AI seems like a really important lever for avoiding power concentration, so we did a lot of our designing to try to avoid monopolies on AI. Then transparency also seems like a really important lever, so we went really hard on transparency.
So it’s kind of frustrating that people — many of whom haven’t even read our thing — say, “They’re concentrating the power in a global regulator,” or something.
Luisa Rodriguez: Right, right.
Daniel Kokotajlo: What else? There’s probably lots of other misconceptions, but I think that’s like the one that bothers me most and sticks out most.
I think there’s a much more innocent one about the pause. Basically this one is innocent because it’s just actually kind of complicated and confusing. In some sense we are advocating for a pause on AI development, but in some sense we are very much not. If you read our scenario, and you read what we are proposing, and what we think would happen if our proposals were implemented, it’s a crazy transformation of society by AI over the course of 10 years. That’s very much not a pause in a bunch of ways.
But the truth is it’s kind of complicated. We are advocating for going slower than you could go at maximum speed. We’re saying don’t do these crazy intelligence explosions. So that’s going slow.
But we are saying you should continue developing AI and deploying it and diffusing it and so forth. We’re saying that, yeah, according to our calculations at least, that’s going to lead to things like GDP doubling every year if you’re doing that.
Then also the actual trajectory that we talk about is more jagged, where there is a literal pause on AI development for like six months in 2029 while they’re getting the verification infrastructure set up. Then it continues at a cautious pace. Then there’s another literal pause in the late 2030s when they run up against the limits of what they can control. Then when they solve the alignment problems, they proceed again. So in some sense there’s two pauses, but they’re temporary.
Which parts of Plan A are recommendations vs predictions? [03:23:02]
Luisa Rodriguez: It’s a bit hard to tell what in the scenario is considered ideal vs a concession to feasibility. How much of each is there in Plan A?
Daniel Kokotajlo: Yeah, I feel a bit bad about this. We had identified this problem before launch and done some things to address it. But we could have been more clear, I guess, and perhaps if we had decided to delay the launch, we could have done more here.
But we have a supplement that talks about it, called “Plan A assumptions.” I think that talks about this question and tries to canvass what’s the recommendation and what’s a prediction.
The high-level thing is everything’s a prediction except for the key recommendations that we talk about, basically. You can go read that supplement and see the things that we consider our main recommendations — those are obviously recommendations, not predictions. Then you should sort of, by default, assume that things are just a prediction about what would happen if our main recommendations were implemented.
That’s the high-level answer. Then there’s a few grey-area cases and things like that we can get into.
Luisa Rodriguez: OK, but to make sure I understand, it’s like you made some recommendations that you think are key to making Plan A go well — everything else is what you think would happen, assuming those recommendations were roughly implemented?
Daniel Kokotajlo: Assuming those things were done, yeah.
Luisa Rodriguez: Are your recommendations basically trying to balance what seems best and what seems possible?
Daniel Kokotajlo: Yeah, basically. I think maybe one way of putting it is we didn’t want to make some recommendations that were basically of the form, “Listen to us and do everything we say forever,” because that’s not politically possible. That’s a bit arrogant.
Instead, we wanted to make recommendations that we could at least imagine being actually done. We talk in the piece about the sort of reasoning and the public discourse and how it evolves and why it makes things like Plan A and Plan S on the table as things that the politicians might actually go for. We wanted to go for things that were within the realm of possibility potentially, and as serious things. But then other than that, we wanted to pick the actually best ones rather than just—
Luisa Rodriguez: The more likely ones.
Daniel Kokotajlo: Yeah, the more likely ones.
Luisa Rodriguez: OK, and so then where are the grey areas?
Daniel Kokotajlo: Too many to go over. But I could give an example.
Luisa Rodriguez: Sure.
Daniel Kokotajlo: So we talk about the citizens’ dividend, and we talk about how it starts off with a dividend for US citizens, but then they extend it as a sort of foreign aid to all human beings. But then they give less dividend to foreigners than they do to US citizens. That’s more of a prediction than a recommendation.
Luisa Rodriguez: Right.
Daniel Kokotajlo: But it’s kind of a little bit of a grey area because we obviously think it’s good to have a citizens’ dividend and we think it’s also good for there to be foreign aid. But is that exact ratio of dividend to foreign aid what we recommend? No, we would want there to be more foreign aid than that, especially in the long run.
I think in the long run we want it to be just actually equal. But that was sort of a concession to reality in some sense. We asked ourselves, “Our recommendation is to do a citizens’ dividend with some foreign aid component,” and then it’s like, “Realistically, how much foreign aid component would probably happen supposing that they did something like this?” Probably they would give less to the foreigners than to the US citizens. So I guess that’s what we’ll write. You see what I’m saying?
That’s an example of a sort of grey area where it’s like there’s parts of it that are a recommendation, but not all of it is our recommendation. If we were in charge, we would do something somewhat different.
Plan A’s likeliest failure mode [03:26:52]
Luisa Rodriguez: If you picture Plan A failing, what do you think is the most likely chain of events that causes it to fail and then follows from the failing?
Daniel Kokotajlo: We talk about this a bunch in the piece. The most likely way that we think Plan A could fail after having been implemented is that the regulators of the various AI industries do a bad job and approve the creation and deployment of AIs that are in fact dangerous, but they wrongly think that’s not dangerous.
Luisa Rodriguez: Right. At what point is this? Is this quite a few years in?
Daniel Kokotajlo: It could happen at any time. It’s most likely to happen relatively early. I think that the longer that the deal has been in operation, the more time the scientific community has to grapple with the situation and the more time the regulators have to skill up, especially thanks to the transparency.
Basically, I’m most especially worried about this failure mode happening relatively early into the deal. We have a little scenario branch that you can go read of what it might look like for this to happen.
The second most concerning failure mode I think would be the deal breaking down. Basically, there’s going to be a lot of yelling. We are realists about this. We’re trying to be realistic about it. We are not proposing a single global authority for AI development. Some people mistakenly think that’s what we’re proposing. But if you read our thing, that’s not what we’re proposing.
Instead, we’re proposing that each country regulates its own AI industry, but that because of the transparency they can see who’s doing what and who is regulating what. If people have a problem with what someone else is doing, they can immediately see it and then they can talk about it and then they can yell at each other, bargain, threaten, plead, and try to get them to stop doing the thing that’s scaring them.
But this is going to be a messy process. Hopefully, eventually it would evolve into a more formalised process that’s more efficient and has lots of technocratic experts making judgement calls.
But at least at first we wanted to be more realpolitik about it and basically just be like: the fundamental thing that the countries have agreed on is the transparency so they can see what’s happening, but then beyond that, they’re just taking things on a case-by-case basis and arguing about what’s fine and what’s not fine. They’re each doing their own regulation, but then they’re trying to adjust their regulation in response to what other countries want them to do and in response to what other countries are in fact doing.
So anyhow, that could go wrong. It could be that tensions get too high and they just really can’t agree on things, or maybe there’s some other thing going on that causes tensions to be high. Maybe there’s a war over Taiwan, for example, that wasn’t caused by AI but is happening. Then as a side effect of the war, they stop doing all this transparency about their AI programmes.
There’s a whole host of reasons why the deal could break down and why they could stop being transparent with each other. Then if they stop being transparent with each other, they’re going to be afraid that they’re going to be racing to superintelligence again, which means they’re probably going to start racing to superintelligence again, which means now we’re in the AI race situation again. Except it’s probably going even faster because they have more compute.
Which means that probably they would destroy the compute because that’s one of the principles of the deal that I mentioned. So that would be a whole messy situation because of the compute-destroyability thing.
We think it would be at least not worse than if they hadn’t made the deal in the first place. And for a variety of reasons, maybe somewhat better. For example, the amount of science and general understanding about AI would have advanced in the intervening years, so we’d be better off from an alignment perspective than we would if we just hadn’t done the deal in the first place.
And in general, more people would have woken up to the effects of AI and would be more prepared, but it would still be quite messy and quite bad if the deal broke down and we started racing again.
What the US can do now to make Plan A possible [03:31:16]
Luisa Rodriguez: OK, I want to move on and spend a few minutes talking about concrete, technical, institutional work that needs to happen in 2026, 2027, to make Plan A more possible.
You’ve already talked about some things that the US could do domestically that would be good for slowing down AI progress in a way that will make safety easier. But it feels like that’s maybe a few steps away from where we are.
What are the literal next steps that you’d like to see the US government do, without any international agreement, to make something like Plan A more likely later?
Daniel Kokotajlo: My answer to this is in the scenario in 2027, our incremental AI policy wishlist.
I think that the limit to AI R&D budgets thing is somewhat ambitious, but I think it’s within the realm of possibility actually. I think it’s more feasible than I think people in DC would expect. I actually think there’s some interest among the researchers at the AI companies to do something like this.
Luisa Rodriguez: Wow.
Daniel Kokotajlo: So I actually think that we could just get started on that immediately.
Other things. Either enforce or repeal the export controls. If we’re going to have export controls, that should be enforced.
AI compute tracking seems good to tell the intelligence community that this is a priority and that they should be trying to find out where the chips are, and see if there’s any covert projects being assembled.
I think that, in general, improving the government AI capacity is obviously very important.
Luisa Rodriguez: Yeah, what does that look like?
Daniel Kokotajlo: The government should be recruiting AI experts and forming agencies within the government that can understand AI and can run evaluations on models and can make safety cases and evaluate safety cases and things like that, can make forecasts about where all this is headed. Yeah, that seems really important.
I think also just transparency more generally. For example, there could be requirements for whistleblower protections. There could be requirements that companies publish model specs or constitutions, and otherwise give more information about how they’re training their AIs to the public — and then even more information to government auditors, so that governments can check that they’re not trying to put any secret agendas into their AIs, for example, or hidden biases. Yeah, things like this.
I think actually this is just scratching the surface. I think there’s a huge list of things like this. Then for each thing like this, there’s a huge list of more specific, concrete things that could be done.
Luisa Rodriguez: Do you feel like that list is written down?
Daniel Kokotajlo: There are some lists like this. I think we have a blog post or two about this. Then of course on our website we say some things, but one of the things we’ll probably do in the next few weeks or months is write up more ideas like this and publish them.
Then there’s other people besides us who’ve also been pushing and advocating for things.
Luisa Rodriguez: So Plan A requires verification technology that doesn’t yet exist at scale—
Daniel Kokotajlo: That’s not true.
Luisa Rodriguez: OK, say more.
Daniel Kokotajlo: I wouldn’t say it requires that technology. I think it’s much less costly if you have the technology.
The way I would put it is: if we wanted to implement Plan A right now, the US and China would say, “OK, we’re going to send physical humans to all the data centres to put their hands on the GPUs and verify that they are cold and off.” That’s something we can do today. We can unplug the machines and then verify that the machines are where they’re supposed to be and that they are off. No technology required for that.
Problem with that, of course, is it’s very costly. It means that all this economic value is not happening because the GPUs are off instead of serving customers.
But you could do it if you wanted to get that going, you could do it today and then you could immediately start creating the new data centres that are going to be more transparent and that have the monitoring devices on them to publish the activity to the internet. You could start building that today and have the existing data centres just off while you were getting that set up. It would probably take, with some sort of crash programme, six months to 18 months to get all that new stuff working.
Then you could proceed with AI development again in the new transparent way, with the new transparent destructible data centres. You could get started right now, but it would be costly because of that.
It would be nice to have built already the monitoring devices and the inference-only retrofitting kits, so that you could allow the current data centres to keep operating and serving customers, and just quickly retrofit them so that they can’t do big training runs — without really interrupting their operation. Then you build the new data centres that do the training. That’s what happens in our scenario.
In fact, if you had even more foresight than that, you could do this without any disruption. You could just make this a requirement for new data centre construction, that they be compliant with the new system. Then after a few years it would just be the way that data centres were by default.
Luisa Rodriguez: If the government wanted to do either of those two things — either do it with lots of foresight or just invest in the technology — what concretely would they need to do and who would be doing it, and how much would it cost?
Daniel Kokotajlo: Obviously we’re uncertain about this, yada yada yada, but our estimate is that it would be single-digit billions to get all the initial hardware developed and manufactured.
For example, the inference-only retrofitting. That gets you off the ground. It means now you’ve started off intent on doing Plan A. Then on an ongoing basis, the new data centres that you’re constructing and making totally transparent, maybe it costs something like 1% or 0.1% more for each new data centre compared to their default cost. So it’s a cost, but it’s well worth it, I think.
Luisa Rodriguez: Right. Who should be thinking about this? What are the steps to actually making sure this happens?
Daniel Kokotajlo: I would say that people with the relevant technical skills, people who understand hardware, for example, and in some cases software, should be trying to build these devices and make prototypes. And companies should be throwing money at this and spinning up divisions to make inference-only retrofit kits, and make different types of chips that have these properties.
Then governments, of course, should just be encouraging this sort of thing. Either by throwing funding at it, like grants, or by basically just saying, “Hey, we want to be doing something like this in the future. There’s a chance that we might require this of data centres in the future.” Just saying that. I think if it was said by the government that might encourage some companies to allocate some resources to it.
I think there’s another thing which is fancier — which I don’t think we talk about as much because we have total research transparency, but which could be really valuable if you’re not going to do research transparency — which is privacy-preserving auditing.
Imagine a situation where all the activity on a US company’s data centre is visible internally. The company can see what’s going on in that activity. Then Chinese auditors show up with a device on which there are some Chinese AIs, and then they plug in and crawl around over all the activity and they look at it all and then they report back: are the rules being followed or is there a violation here? Then they’re deleted and the device is destroyed, so they weren’t able to exfiltrate any secrets. All they were able to do is just, “Yes or no, are rules being violated?” Then, of course, we do the same thing over to China.
In order to have that sort of setup, you need to have a fancy piece of technology that doesn’t really exist yet — but maybe could exist if we built it up. That could be really valuable because it would allow us to do this sort of auditing and get exactly the information that we want, without any more information than that leaking, if that makes sense.
Luisa Rodriguez: Cool, yeah, yeah.
Daniel Kokotajlo: But someone needs to build all of that, and derisk all of it.
Luisa Rodriguez: The scenario assumes that labs operate at Security Level 5, which is nation-state-resistant cybersecurity. Right now they don’t. What needs to happen for labs to get there?
Daniel Kokotajlo: Oh yeah, that’s another thing. Previously I mentioned that, in some ways, the total research transparency is a gift to China because it’s sharing the algorithms directly with them.
Well, they’re probably getting the algorithms anyway because security is not very good right now. It’s not even that big of a concession at the moment. But obviously we think more security is better.
You asked what is the pathway to get there?
Luisa Rodriguez: Yeah.
Daniel Kokotajlo: Well, that’s one of the things that comes along with the new data centres. If you’re going to be serious about this sort of thing — and you’re requiring that there be new data centres that are built in a transparent way — in addition to the transparency requirements that we think the new data centres should have, you can also add on security requirements to them. You can make it so that it’s extremely difficult, in fact impossible, for even a nation state to exfiltrate the weights, for example.
One mechanism for this is just having a bandwidth limit, so that it’s not even possible for the weights to leave the data centre through the only cable through which information can leave and exit the data centre — because the weights are too big to leave through that cable. That’s an example of something you could do. But there’s a whole bunch of other best practices that you should totally do as well.
Again, in our Plan A scenario, they first do a temporary pause where they stop all new training runs and they refit the existing data centres to be inference only while they build the new data centres that are going to be much more secure and also much more transparent and also in these locations where they’re destroyable and so forth. And that takes time. But with a crash programme, we think it can be done in six months to a year, or something like that.
Luisa Rodriguez: OK, so is it basically the case that if we took a bunch of steps, we already know the steps that would be required, and if we implemented them we’d be there?
Daniel Kokotajlo: Basically, I think.
I think part of what happens in our scenario is that they’re doing things last minute. They had prepped some of these things in advance, but if they had decided to implement Plan A in 2027 instead of in 2029, then the process would have been much more smooth. Naturally the data centres being built in 2029 are built to the new code, so that’s just how it’s going.
Luisa Rodriguez: OK, let’s leave that.
How AI 2027 is holding up [03:43:05]
Luisa Rodriguez: I just have one more question for you. Relative to your expectations from one to two years ago, how do you think things are going? I guess alignment work, how seriously various governments take AI risk, just a broad range of things.
Daniel Kokotajlo: Unfortunately, things are going roughly as I expected. You can still read AI 2027 and it still seems like, yeah, we’re sort of going down that path.
I used to say that the alignment situation was better than I expected, and the governance situation was worse than I expected. But that was what I would have said a year ago or two years ago, but now I almost say the opposite. Compared to a year or two years ago, I would say that the governance situation is better than expected and the alignment situation is a bit worse than expected.
In particular, the Hugging Face rogue AI incident is a more egregious example of misalignment than I expected to be happening at this time. You can tell by reading AI 2027, for example, where we talked about the misalignment over time and nothing this egregious happens in 2026. That’s, I guess, a minor example of things being worse than I expected on the alignment front.
Then on the governance front, I think that the silver lining of all the battles between Anthropic and the Trump administration is that the Trump administration is not being bowled over and captured by the leading AI company in the way that happened in AI 2027. They might. We’ll see what happens. Maybe it’s partly a personality thing, and maybe if OpenAI was in the lead then they would be.
But at least the way it’s currently going is that it seems like the administration is more willing to bring the foot down on the companies than I expected, for better or for worse. But since the situation seems pretty bad to me, it means I still have some hope that they’ll do it in the good way. Whereas previously I was expecting pretty bad things, and now it’s like I have a little bit more hope that they’ll do the good things.
Then also the broader public is just gradually starting to take all this stuff more seriously. Various senators and congressmen are talking about loss of control risk and so forth, but overall things are not that different from what I expected. These are just slight changes.
Luisa Rodriguez: Is there anything we haven’t talked about that you want people to know?
Daniel Kokotajlo: Yeah, I think I want to leave people with this high-level point about what we’re doing and why. We don’t want this to be the end of the conversation. It’s more like the beginning of the conversation.
We are not confident that Plan A is the best plan. We see a lot of problems with Plan A, and a lot of ways it could go wrong. We just think it’s the least bad plan that we’re currently aware of. We think that the alternatives that other people — including the major AI companies — are proposing seem dramatically worse in a number of ways than Plan A.
We’re hopeful that as people wake up to what’s coming and take it more seriously and start gaming things out, that people will think about all these plans — including Plan A — and take the best elements of them and combine them. We’re hopeful that what ends up happening in practice will be better than Plan A.
That said, what we actually expect is that what ends up happening in practice will be worse than Plan A.
Luisa Rodriguez: My guest today has been Daniel Kokotajlo. Thank you so much.
Daniel Kokotajlo: Thank you. Thank you for having me.
Our podcast team is hiring [03:46:45]
Zershaaneh Qureshi: Hey listeners! If you’re enjoying this conversation, then I’ve got to tell you we’re actually hiring people to help us make more episodes like it.
We’ve got three open roles on our team:
- A producer role
- A production coordinator
- A special projects role
These roles basically range from shaping the content of episodes to running the production pipeline to driving forward new projects independently. You can find more details at 80000hours.org — just head over to the site, click on “Work with us.” Just bear in mind that applications close on the 30th of August, 2026.
Related episodes
About the show
The 80,000 Hours Podcast features unusually in-depth conversations about the world's most pressing problems and how you can use your career to solve them. We invite guests pursuing a wide range of career paths — from academics and activists to entrepreneurs and policymakers — to analyse the case for and against working on different issues and which approaches are best for solving them.
Get in touch with feedback or guest suggestions by emailing [email protected].
Our crash course on transformative AI
We've carefully selected 10 key episodes to help listeners get to grips with the potential upsides and downsides of powerful, transformative AI.
Check out 'The 80,000 Hours Podcast on AI'
Listen here, or anywhere you get podcasts:
If you're new, see the podcast homepage for ideas on where to start, or browse our full episode archive.







