Transcript
Toby Ord is back — for the 5th time! [00:00:00]
Rob Wiblin: Today I’m speaking with Toby Ord, senior researcher at Oxford University’s AI Governance Initiative, and the author of The Precipice: Existential Risk and the Future of Humanity.
Welcome back to the show, Toby. I think it’s your fifth appearance.
Toby Ord: It’s great to be back.
AI self-improvement might not matter [00:00:14]
Rob Wiblin: It feels like today everyone is talking about recursive self-improvement (RSI), or at least like everyone in my circles is talking about AI recursive self-improvement.
Do you think that RSI will lead to an intelligence explosion, like a massive takeoff in AI capabilities?
Toby Ord: No.
Rob Wiblin: Go on.
Toby Ord: [laughs] At least I think that’s unlikely. However, the chance that it might happen I think is credible. And this possibility really is one of the biggest issues in AI at the moment.
Rob Wiblin: So why do you think it probably won’t work or won’t pack much of a punch?
Toby Ord: There’s a couple of reasons. One is that I think that AI research involves a lot more than just programming, and a lot more than just running the kinds of simple experiments that have been automated so far. So I think that there’s a pretty reasonable chance that some major breakthroughs are still needed and that AI assistance isn’t going to be able to provide that.
Rob Wiblin: What if AIs do get better, such that they can more comprehensively replace the work that staff at the companies are doing?
Toby Ord: Well, there’s a lot of different types of work at a company — this includes, say, cleaning and catering, also programming and other aspects of software engineering, and then running some simple experiments.
But it also includes, I think, genuine strategic decision making and also AI research. That is not the case of programming — where you know what you want and then you type it in — but a case where you don’t know what you want, you don’t know what changes to these basic architectures will lead to solving some of these long-standing issues and really letting these systems take off.
And there’s not much evidence at the moment that AI systems can do that kind of thing. They might be able to really speed up the programming. But even if they made programming instantaneous, I think that may not actually change the timelines all that much. So they really would need to be able to help with these harder aspects of research.
Rob Wiblin: Is that the crux of the issue? If they could replace human researchers, like the full suite of things that the staff are doing, do you think then we would get an intelligence explosion?
Toby Ord: I think it’s quite plausible. There’s still some other issues. When you think of an intelligence explosion, you tend to think that this rate of progress over time is kind of bending upwards in some way, but most people think that then it will eventually kind of plateau off, and so it will start to bend back down as you reach some kind of limit, or maybe the increasing complexity of the system starts to overwhelm your ability to improve it.
So most people think it will bend back down to flatten off. The real question is, before that, is there a phase where it’s bending upwards? Or is it just the phase that’s flattening off? I don’t think that it’s clear which one of those it is — even if it could replace all of the human skills.
Rob Wiblin: What are the most distinctive aspects of recursive self-improvement that make it different than humans doing the work?
Toby Ord: There’s an interestingly long history of talking about recursive self-improvement — quite a few decades of talking about it, followed by now we’re at the situation where a lot of people are suggesting that they’re doing it.
This started famously with I. J. Good, who talked about this possibility of an ultraintelligent machine. What I think a lot of people forget about his story of recursive self-improvement — although he didn’t use that term, he used the term intelligence explosion — he was thinking of an AI system where, once you have an ultraintelligent machine, one that’s more outside the human range, it’s beyond anything that a human can do, and in fact it’s beyond what the entire research community can do; once you’ve got that point, he said that this is the last invention you need ever make because it would be better at everything by definition.
One of those things is building more intelligent AI systems. He thought that this would lead to this ultraintelligent machine building ultraintelligent machine number two, which then builds number three and so on — and that this would have an explosion, leaving humanity far behind, although perhaps it would also peter out at some very high level. That was the first idea.
Then there’s a big change to it by Eliezer Yudkowsky around about the year 2000, where what he considered was, maybe you don’t need it to be ultraintelligent to start this off.
In I. J. Good’s version, the AI community have had to get all the way up to ultraintelligent machines without any recursive self-improvement. But in Yudkowsky’s version, he thought, well, maybe if the one thing the system was good at was improving its own code, then you could have a system that’s below the human level and easier to build, but then it keeps improving itself and improving itself. And now that it’s so good at improving its own code, it works out how to improve other aspects of itself and then can ultimately blossom into this full set of capabilities at some very advanced level.
But his version still involved this idea that it was a bit like ‘AI improver’ is, say, your programme written in Lisp or some other kind of symbolic system. And then you run ‘AI improver’ on the source code of ‘AI improver,’ and you put that in a loop, and then you hit return — and then this thing whirs away and improves itself like a thousand times, say over the next day. And then you’ve got this superintelligent system with no real opportunity for humans to be in the picture anymore.
They’re the previous versions. Then the modern conception that the labs are talking about now, it’s more like you’ve got companies with thousands of human researchers and programmers, and also they’re building AI. And so far the AI’s abilities depend upon what these researchers do. But we could have a little bit of input from the AI itself where it kind of feeds back in this kind of feedback loop.
For example, at DeepMind they developed an AI system that found a slightly faster matrix multiply algorithm. Then the new chips for the next generation of AI could be, I don’t know, like one part in a thousand faster or something. It wasn’t that big a deal, but it was interesting. That was a first stage where the AI improved its successor. And there have been a few other things like this. The most notable these days is the AI coding systems that could take on a whole lot of the coding responsibilities for creating the next AI.
But we’ve got a system there where there’s humans and AI together, and the AI is not way below the human level, it’s not way above, it’s somewhere on the lower outskirts of the human level at a lot of things. And it begins by inputting a small amount into the process. But that amount could increase and increase until it’s mainly driving the process.
Rob Wiblin: Yeah, I think the companies imagine it now as a gradual process of replacing their staff, where gradually the AI is going to be able to do more and more of all of the things that the humans were doing. But I guess at the moment the AIs and the humans are incredibly closely integrated, they’re working together basically as colleagues, because there’s lots of things that one can do that the other can’t, and vice versa. What implications does that have that maybe shift people’s expectations relative to…
I think so much of this conversation has been informed in a way that people don’t always appreciate by the Eliezer Yudkowsky vision that has been so dominant in shaping the stories that people tell about how it’s going to play out. But if it’s not playing out that way, why is it different?
Toby Ord: I think the biggest change is one with respect to safety, which is that there’s much more of an opportunity for humans to remain in the loop.
I remember noticing back, I don’t know, 15 years ago with the Yudkowsky story, that he had this compelling vision about why AI could be really dangerous. And at some point I noticed that so much of the danger was coming from this recursive self-improvement concept. And I remember thinking, maybe there should be a moratorium on this about 15 years ago. But it was outside the Overton window to talk about not doing this thing that no one could do. So it’s a little bit hard to get these things off the ground when you’re too early.
But yeah, it was driving a lot of the risk because it was this thing where you had to get everything absolutely right. You’d had to understand human values and be motivated by them at the point where you start this process. And he was imagining a system that can’t even speak English at the point where you start the process. So how do you write something—
Rob Wiblin: Even begin to communicate, yeah.
Toby Ord: Maybe you write this English text that, once it’s halfway up, it then can read it and then it tries to act on that and so on — but really dangerous without being able to have any way to kind of interact with the system once it’s going, no ability to learn from feedback for us.
Whereas this new version is substantially more gradual, but also it has a lot more opportunities to understand what’s happening. We even have interpretability tools to try to understand what’s going on inside the systems while they’re climbing up this approach. So I think that it does look somewhat safer than these old versions.
Rob Wiblin: Something that I find quite frustrating is it feels like, I guess I and my colleagues have been thinking about this for years now, the recursive self-improvement possibility. It feels like we’ve gotten almost no greater clarity on how fast it’s going to go and at what level it will plateau.
Did you also have that view that we don’t know whether it will work, we don’t know how long it will last, and we don’t really know on what level it’s going to cut out or peter out if it does work?
Toby Ord: Yeah, that’s right. I’m actually working on a paper at the moment trying to understand these possibilities of really explosive growth, which I do think is credible and I’m concerned about it.
But a lot of that is trying to understand the dynamics. Is this the kind of explosive growth that’s exponential? Or is it the kind of thing that is faster than an exponential, such that it would have this kind of vertical asymptote — where there’s a particular time at which the level of this intelligence goes to infinity? That kind of dynamic, as opposed to the more typical compounding exponential growth from populations and things.
Some of the debate is over that. But even when you’re debating that kind of question, almost everyone in this area thinks that the intelligence is going to cap out at some particular maximum level anyway.
So to some extent the exact shape of it may end up being moot, if it caps out not that much above where it currently is, for example. Or there’s these questions of: sure, it’s an exponential, but is the doubling time on the exponential like a week, or is it a decade? That’s a big deal. And often the theoretical analysis doesn’t shed much light on that.
One of my colleagues, Tom Davidson, has great work on this, great theoretical work where he’s looking at all these questions about the theory of it, and will it have this vertical asymptote which gets called a ‘mathematical singularity’ — not to be confused with the more sci-fi concept of the singularity, although maybe they could be related.
But then his bottom line is saying something like he thinks that it’s most likely that there’ll be something like five years of AI-research progress that happens within one calendar year and then it starts to plateau, something like that. And it’s like, does all of this stuff go to infinity, or do this thing in the limit as time goes on forever, what happens to it? Then it comes down to some little rule of thumb, like it could go five times faster for a bit or something.
Rob Wiblin: It would be quite a hectic year, I suppose, if we had like five years of past AI progress compressed into one. But I suppose it doesn’t feel like it’s impossible that we would be able to manage that.
Toby Ord: No, and it also depends on where that leaves us. Suppose we had the previous five years compressed into one year. That would have been very disruptive, all from pre-ChatGPT through to where we are now. But it would leave us at a point with a somewhat manageable technology.
So it’s not clear exactly where this thing does plateau. It might plateau at a point that is not vastly superhuman and that is controllable. Or it might be that the five years takes us through that period where things get out of control. I think that would be a big deal. Obviously, the five years itself is just guesswork.
Four ways AI self-improvement is dangerous [00:12:38]
Rob Wiblin: You said that you thought in the picture that people have had historically, much of the risk came from recursive self-improvement. Why is that distinctively risky? And do those types of risks carry forward to the kind of recursive self-improvement that we might see in the next few years?
Toby Ord: If I cast my mind back to how this is being thought about before large language models, part of it was that there was this kind of question that got called ‘the value loading question‘ — where if you take the richness of human values and all the things we care about, and the way that if you miss out important aspects of the human condition, that it could lead to this impoverished future that is dystopian in some way.
Well, Stuart Russell was one of the few people who predicted what happened, which was that the AI system has read all of the books we’ve ever written, which are millions of pages of text. On almost all pages, some human is judging some other human about something, whether it’s Lizzy Bennet judging Mr Darcy, or what have you. So you can learn vast amounts about this. And that was outside of the previous conception because it was going to happen so quickly that it wasn’t clear there was any moment where you could get it to understand what we care about.
And also that’s one aspect, but maybe the broader aspect is that kind of feedback, some ability to actually say, if you’ve got some experiment, “the experiment’s going off the rails — quickly, pull the emergency stop button” or something like that, or to learn other aspects about how it’s progressing and to then be attentive to them.
Maybe if you’re planting a tree in your garden that you’ve never planted before, if you can go out every day and check on it and see that its leaves are wilting a little bit and adjust and so on, it’s a lot easier to make that tree survive than if you had to build a machine that will appropriately measure out water for it every day or something, and just turn that machine on and hope the whole thing works.
Rob Wiblin: So that’s the risk that stood out to people 15 years ago, I guess. What’s the picture now?
Toby Ord: Yeah, so now there’s a few different kinds of risks that come from this. I think the fundamental thing is that everything’s going a lot faster. That’s where a lot of the risk would come from.
Ultimately, speed itself is not quite the risk. Although if it did bring forward some apocalypse by a few years, then all our lives would be shortened by a few years, which would actually be a pretty big deal. But it’s probably even bigger than that because probably what would happen is it would change the probability that things go off the rails and that things would get a bad outcome.
And the way it could do that through sheer speed is that it could change the speed of the AI activities that are increasing risk, such as improving the capabilities of these systems faster. It could make more of an impact on that than it makes on some of the processes that are reducing the risk, such as technical research on AI alignment or other aspects like AI governance work. I don’t think it’s going to speed those things up as much as it would speed up the capabilities work of improving its own capabilities.
So if these things get out of whack, then what you’d expect is that for any particular level of capable AI systems, that our ability to control them is lagging behind. So we’d pass through all of these thresholds less prepared than we otherwise would be.
Rob Wiblin: So is recursive self-improvement itself risky, above and beyond the fact that it speeds things up? I guess I feel quite unsure about this. And it seems like people disagree at the companies and outside of them.
Imagine that it didn’t speed things up. I guess the obvious way that it’s more dangerous is that the AIs are doing the work themselves, and so it’s more difficult to monitor someone else doing something than to be doing it yourself. Is there more to it than that?
Toby Ord: Let’s call that the second one, the second source of risk. If you imagine producing one AI system, let’s call it GPT-7 and then that produces GPT-8, which produces GPT-9 and 10, and so on. If any of those in that line were to be misaligned, and to be able to scheme about what they want to achieve, then they would have an incentive to misalign the systems that they were building and to make sure that whatever its peculiar objectives were that differ from human objectives, that that was maintained in the future systems. That’s the second big source of risk, which is also a really big deal independently from the first.
And it could have different remedies for it. If you look at each of these different ways that recursive self-improvement could be risky, there could be different forms of governance or different kinds of rules applied at the labs in order to manage them.
So then a third one would be that, when humans create a new technology — say guns: first we create muskets, and then we make rifles, and then eventually we’ve got handguns and machine guns and things like this. And we’ve got this opportunity though to learn from the intermediate stages. So we don’t unleash the most powerful form of the technology on society all at once. Or with cars, we had much slower cars before they could go at the current speeds, so that gave us opportunities to learn. Also there were only a few of them around at first, instead of everyone having a car.
Having those kinds of intermediate levels is extremely useful for society. We’re actually really quite bad at predicting what’s going to happen if some new technology were to arrive and to regulate it in advance. But if we get to witness some of the ill effects on a smaller scale, then we can learn from that.
In this case, if things are going — let’s say they’re going five times faster — maybe instead of it being the model that would have been released in 2027, in 2027 we get the model that would have been released in 2032 or something. So we have some really big jump up there. If we get that, we don’t get the opportunity to learn from these intermediate things. And it could be that there’s one that really would have created a lot of havoc, but at a level that was ultimately manageable by society — but we really learned our lesson. And in this case we might not get to learn that lesson.
Rob Wiblin: Yeah, is there a fourth one?
Toby Ord: I think a fourth area is that, if you imagine the rate of progress just with human-only research in AI — let’s say that’s kind of ticking along — and then imagine that there’s a leading lab where their research is improving the capabilities over time. And then there’s a trailing lab that’s a year behind them.
The difference in the capabilities of the leading lab and the trailing lab, it’s noticeable, but it’s only so big.
Whereas if you have this recursive self-improvement and the capabilities really bend upwards steeply at some point, then the first lab to go through that process, it could be that — if we go with Tom Davidson’s rule of thumb of five years of progress in one year — it’ll be then five years ahead on the old scheme, where the lab that was a year behind is. So that is more likely to lead to a winner-take-all scenario, which encourages racing, and also to lead to a final outcome where there’s much more concentration of power towards whoever gets there first.
I think that’s a fourth reason. There’s probably more. I’m hoping that people will take up this idea of actually trying to think about this, and trying to work out what are the different sources of danger, and then what kind of remedies would be required to deal with them, how much danger are they producing, and so on.
A US-China treaty on superintelligence is possible [00:20:46]
Rob Wiblin: I think people who are very sceptical that there will be any slowdown, the argument that they most lean on is just the competition between the US and China and the distrust between the countries will make it impractical to delay the arrival of things. It’s going to, if anything, inspire the government to want to push forward to superintelligence as quickly as possible in order to avoid being vulnerable to a country that they see as partially an adversary.
Do you think an agreement between the US and China is more likely than those folks believe?
Toby Ord: Yeah, I think, at least in the long run, I think it is quite credible. In the short run, could there be a deal between China and the US on recursive self-improvement, and how would that work? I haven’t thought about that one too much, but that could be a bit harder.
It does depend on the nature of the deal. But if the idea was something more like a moratorium on superintelligence — that we’re willing to go a certain distance into the world of AI and AGI, but that systems that completely dwarf the intellectual powers of humans are off limits — I think that such a deal would be very much in the interests of both countries from the perspective of AI risk, that once they realise that there’s a very credible chance that by allowing these things that they’ll be killed and that their children will be killed and their whole culture erased from the Earth, perhaps — if that is what the evidence suggests, then if people ultimately wake up to the evidence, I think it’s just in their own interests generally, if they can, to negotiate some way to not go there.
Rob Wiblin: So the game theory is a lot easier if both governments basically are extremely scared of superintelligence, and they think the risk posed by the AI itself, or the risk posed by AI advances, is large relative to the risk posed by the adversary country that they’re worried about falling behind against.
But I’m not necessarily sure that the countries will end up being as scared of AI as that, to make such an agreement easy.
Toby Ord: I think that it’s pretty likely that they will be. At the moment, it’s only people who are looking out quite far ahead who can see some of these possibilities clearly, and are willing to trust some of this extrapolation about what’s going on. That’s only a subset of people.
But I think of it like with COVID, where I remember — partly by reading some of your writing on it at the time — really tracking this early on and being very much ahead of the curve, and really looking at these graphs and working out how many cases there really were and so on. And overall, I would say this just brought me like two weeks ahead of everyone else.
So there were things that I was deciding — for example, to cancel my book tour, going to America for my book — but two weeks later it was just totally obvious. And like anyone would realise that you have to cancel it. In fact, planes might not have even been flying at that point.
And so I think there’s something here, where for us to see now that — in 2030 or something — these systems could be extremely dangerous requires something of a leap of logic or something. But to see that as we approach it will become increasingly obvious.
And so if we approach it gradually enough — such that before they’re at the level where they could take over the world if they wanted to, they’re at a level where they could do other amazing things at a similar level to taking over the world, and then before that they’re just a little bit below that, and so on — then it should be actually pretty obvious to people. You don’t have to look into the future, you just have to look into the present, really.
So I think that we neglect that. We think a rare combination of things to really take ideas seriously is needed. But I think that won’t be needed. And in fact, just like with COVID, just the natural self-preservation instinct of “I don’t want to die, I don’t want my children to die” will kick in, and will be enough actually. So it partly depends on the gradualness.
I don’t know whether that will quite be enough, but it will definitely make it easier and easier over time to see where this is heading, if it really is heading there.
Rob Wiblin: The other pushback I get on this notion is firstly, if you talk to people who have done negotiations with China or been involved in that kind of diplomacy, on the US side they’re unbelievably jaded. They feel like there’s a lot of bad faith on the Chinese side, that it’s very hard to get them to sincerely agree to anything and follow through on it. It wouldn’t shock me if on the Chinese side there’s also some scepticism about the US’s willingness to follow through on treaties.
And another concern is that the national security community, when you describe a new, incredibly dangerous thing, military people, their first impulse is not “we have to ban this”; it’s “we have to have this before anyone else.”
And we see that with the cyber capabilities of Mythos and Fable, that there’s a certain hunger to grab those powers ahead of anyone else. Do you think those are severe barriers?
Toby Ord: Yeah, there’s definitely a barrier. But reading, say, Situational awareness by Leopold Aschenbrenner a couple of years ago when that came out, it painted a very hawkish picture of America maintaining its advantage on AI and using that in order to be able to disarm and subjugate China — which might have seemed very attractive to hawkish people in the administration.
But if you think of it from China’s perspective — suppose that there is a technology that will mean that you can unilaterally just disarm this other superpower, and possibly have regime change and so on, and finally be able to get all of your policy objectives at this international level with them — then China would be in a position where its nuclear arsenal, the value of it goes to zero at the time when that system’s turned on.
If so, there’s a strong incentive to threaten nuclear retaliation for them to even cross lines approaching that level of power over them. So I think that this alternative strategy to having some kind of moratorium — the strategy of getting there first and then lording it over them and using that to disarm them — I think doesn’t really work because of this issue.
It’s actually very similar to the situation when Ronald Reagan was doing this Star Wars programme (a missile defence shield against Soviet missiles), where again, it sounds defensive, but once you have that shield, all of a sudden the value of Soviet nuclear missiles goes to zero. So they’re strongly incentivised to strike before you do that.
And so all of a sudden you’re creating this kind of first-strike incentive. I think that this is not a good strategy actually because it’s failing to take into account what’s the best response to this strategy, which is ultimately threatening nuclear war.
I don’t think that you can actually just seize this technology, at least not if your opponent’s paying attention. You could maybe only do it if they just don’t realise that the option’s available to them or something. But that’s usually when RAND or other senior people give strategic advice, they assume the best response from the opponents. They try to take this game-theoretic approach of: “Assuming that they’re not an idiot, what should you do?”
Whereas I feel that in this case it’s a little bit like moving your queen somewhere where there is a subtle way that the queen could get taken, but you just hope they don’t see it. And that’s not how this advice normally actually works. I think that there’s been a failure to consider the best response. So I think that counter to the idea of “we should get it first” doesn’t really work as well as people think.
Rob Wiblin: I guess especially not when you’re shouting about it in the press in an extremely easy way to pick up.
Toby Ord: Exactly. And it’s almost like: “We’ll move the queen here. It can be taken, but we’re hoping they won’t see it.”
Rob Wiblin: “We’ll write about it in The New York Times, but we’ll hope they won’t read it.”
Toby Ord: Exactly. I mean, I’ve heard that there are five different Chinese translations of that essay, for example. So they’ve definitely heard it.
I think that approach of “we just get there first and then have all the power on the world stage,” I think is overrated. There’s also this question of is it possible to verify it? And so on. How do we trust them?
And the jaded negotiators, it is a problem that they’re feeling like that. But I think that they’re jaded on discussing a whole lot of issues that are a lot less important than this. I would say this issue, to those negotiators, is more than 100 times as important as any issue they’ve ever negotiated before. If the negotiator was told they’ll get killed if they don’t succeed in this negotiation, they’ll be like, “Oh my God!” They won’t be feeling jaded at that point, right? They’ll be feeling, if anything, a little bit too excitable or something.
So because the stakes become really high and it’s very much in their interest to actually reach some kind of agreement, I think that it can push through that barrier. But there are these questions about how would you make sure that they’re holding up their end of the bargain? I think that the key aspect that you would need is basically to get it to be in China and the US’s interests to have such an agreement, if it could be verified. And I think it is. And I think that they will start to see that as we get closer.
Then we also need an ability to make the agreement actually bind them, and then they could police their own spheres of influence. So you might not have to get everyone else on board at first. So the key step, I guess, is the verification.
Rob Wiblin: Yeah. How would you do the verification, do you think?
Toby Ord: I think there’s a bunch of ways to do it, and it depends on exactly what the technology is at the time and so on.
People have often been thinking about some very technologically advanced verification mechanisms, such as having the next version of the NVIDIA chips be able to know exactly where they are in the world with GPS. And if someone blocks the GPS, then maybe the chip shuts down. It will only start working again if it can share where it is, and so on. I mean, that would be great if we get those things. Then, for example, there would be credible ways to prove that your chips aren’t actually working on advancing AI further.
But even I think people forget that the universe of possible ways of doing verification is very large. And I think the Cold War is really illustrative. For example, one of the key things was that the US and the USSR wanted to reduce the number of nuclear-capable bombers that could threaten the other person’s home soil.
Eventually some bright person, thinking outside the box, worked out that these bombers were on the runway, and they could saw them in half with a huge circular saw. And then, with tractors, pull the halves apart from each other so that spy planes from the opposing side could witness this and credibly see the destructive capacity be destroyed. As far as I understand, those planes are still on the runways, in halves, because someone was like, “Hang on, this is a way you could do it.”
That’s the kind of thing you come up with after you’ve been thinking about this for quite a long time. Whereas, at the moment, there’s a lot of people who’ve thought about it for less than one day, like less than eight hours full-time thinking whether verification is possible, who say, “It’s probably impossible, and so I guess we will all die.” I feel that is a bit of a ridiculous approach.
It’s reasonable to think it might be impossible to do the verification. Maybe we’ll think about it for years and we won’t find any ways. But I think that there are a lot of credible ways, and that’s an example that you couldn’t have initially done. It required the Overton window to move a bit before that one was possible — for the US to say, this will literally involve destroying capacity we’ve already built — and that would have been annoying at first, but once they realised that it would also destroy the same amount of capacity on the other side, you start to think even though we’ve already built it, it’s probably worth doing.
I think an example that would be very analogous to that is: at the moment, there’s a lot of chips that aren’t being tracked, and where we don’t know exactly how many have been smuggled into China and things like that, but you could have, say, like-for-like destruction of that capacity.
As a simple example, you could say the US agrees to bring a million GPUs, or a million H100 equivalents’ worth of GPUs, to Geneva, and China will bring a million there as well. And what they’ll do is they’ll both test each other’s GPUs to make sure they all work, and then they’ll both destroy them. That would be the simple version, but probably you could do better than that and just both put them into escrow. For example, that compute is now being used for some kind of humanitarian purpose by an international organisation, that they both have people on the ground.
Rob Wiblin: And anyone can inspect what’s going on.
Toby Ord: Exactly. But there are ways like that, where even if you don’t know exactly how many chips they have and so on, you could at the very least do one-for-one kind of swap offs like this.
And that’s the kind of thing you come up with one person just thinking for a while about it. I’m sure there’s much smarter approaches you could do as well. I do think that we just haven’t spent much time at all thinking about ways to verify, once the Overton window is broad enough.
Rob Wiblin: Once people are really motivated.
Toby Ord: Yeah. So that idea: I think now we’re at the stage where at least maybe some of the people listening to this will think, “Oh yeah, I guess you could do that.”
And maybe some other people will be thinking, it’s insane to destroy or give away these GPUs to this neutral party. But I think it’s the kind of thing that you can now think of as like, “Yeah, I guess you could do that,” whereas five or 10 years ago you kind of couldn’t even really think about that.
And as we get closer, I think just more and more things will start to seem like, “Hang on, if the alternative is that my children get killed, yeah, I guess I’m prepared to do this thing. It didn’t sound great at first, but now that you point out seriously bad alternatives, is it better than that?” And that the answer will be that, yeah, it’s way better than that.
Rob Wiblin: Yeah. I feel like the distinctive Toby Ord perspective over the last couple of years has been: “What if society does realise or come to think this is a very serious issue, it’s a very big threat, and what if they’re not idiots about it? And what if they do the sensible thing?”
It’s surprising that that has been quite a contrarian take, or quite a contrarian expectation, and not many people are imagining how that would play out. I guess it also requires that things be a bit gradual and not really explosive.
Toby Ord: This is another thing that goes back to that Yudkowsky kind of idea of an intelligence explosion that’s completely automated and happens really quickly. So because of that, he was imagining — when he first started ringing the alarm bells about AI risk — that it would take you from a system that was only of academic interest to a system that could take over the world if it wanted to, without humanity waking up to that outside of some narrow part of the research community.
Whereas what we’ve had is something much more steady, where the intermediate phases are visible. I guess it’s related to this idea that humanity is actually OK at learning from feedback, once we try things and we find out that they’re bad for us, we can adjust our course. So because things have ended up being more continuous than that, and perhaps they will remain so, you can then anticipate that these opportunities will arise.
Whereas recursive self-improvement, one of the problems is it may mean that we go from systems that are quite impressive but not amazing to systems that are all-powerful without that being visible. If so, then we’re back in the world—
Rob Wiblin: Back in the difficult world.
Toby Ord: Yeah, but I do think that a lot of people aren’t thinking quite seriously enough about how easy it will be to notice these things later on.
Could we ban superintelligence? [00:37:06]
Rob Wiblin: If we’re going to have a moratorium on something — and we could imagine like multiple different levels of moratorium; it could start with just a couple of companies agreeing not to do something that they regard as a dangerous research practice, and then perhaps one government can get involved and then you could have an international treaty — what should it be about, what should it be on, and how would you potentially design it?
Toby Ord: Yeah, look, bottom line is I don’t know. I think that there are heaps of options that are all quite interesting. When I think about recursive self-improvement, here are a couple of different things that you could do.
One of them is you could say there’s some level of AI capabilities that we’re not going to just rocket through — that it’s high enough above our current levels that we would stop if we reached it. That’s one approach.
You could also have a speed limit. You could say there’s a certain amount of capability increase per year. Maybe it’s the speed that’s the problem, so you could set a limit that we’re not going to go faster than a certain amount.
There’s also ways that are somewhat arbitrary but also verifiable. For example, companies could say they’re going to do some AI coding, like every company in the world is doing AI coding. It’d be somewhat weird if the AI companies that are making the AI that does the coding are the only companies that are not using it. But there could be a limit on how much compute can be used for that. Say that they will not use more than some very large number — that’s somewhat arbitrary — of tokens on AI for AI R&D automation. So they’re allowed to use a whole lot of compute when training their models in the normal ways and so on, but this approach of “exactly how many tokens are being used per researcher” or something, at the lab.
And you could put a limit on that and people would ask: “Why is it here, and why isn’t it a little bit higher or a little bit lower?” But the point would be that so long as we all agree on some particular limit, then it could actually be fairly easy to check — at least within Western labs — whether they’re all following it or not.
Rob Wiblin: What do you think are the pros and cons of trying to set a speed limit vs trying to just stop progress completely, which is something that at least some people have suggested?
Toby Ord: I guess if the idea is to stop RSI at all, like if it’s some kind of moratorium about recursive self-improvement itself, then I think it’s very difficult to have a thing never get started, because I feel like they’ve already started. So you could say: “OK, you’ve started and the mistake was to let you start. And now we’re going to remove all of that — and no more AIs making better matrix multiplication algorithms, and no more AIs being used for coding of other AIs and so on.” You could try to roll it back, although it’s challenging to do so.
In general, there’s a question about which lines are legislatable and so on. And so that’s one reason why I’ve generally been a bit of a fan of compute thresholds on things, because I think that it’s easier to write them into law and it’s easier to inspect them and so on.
Whereas if you say you’re not allowed to do anything that counts as recursive self-improvement, which is any way that you use AI in order to help with the AI R&D process, there’s probably going to be a lot of corner cases and so on that are quite hard to police. So that’s a less good line to draw. Doesn’t mean it’s impossible, though.
It could also be that you say, yeah, it’s going to be hard to work out. What’s going to happen is that, if we think there’s a violation, it’s going to go to this set of five judges, and if a majority of them rule that it’s a violation, then it’s going to count as a violation. You could just say we’re just going to do that — like how courts actually decide law in the real world.
Rob Wiblin: So one source of benefit from a moratorium would just be slowing down progress so we have more time to learn through trial and error.
I guess a different reason to have a moratorium or some limitations on recursive self-improvement is so that you are more likely to retain a human in the loop who’s monitoring what’s going on and might pick up misalignment or attempted, I guess, sabotage of future models by an AI that was designing its successor.
Do you have any thoughts on how you can encourage that to keep occurring?
Toby Ord: Yeah, I think that there’ll be a lot of incentive actually to keep a human in the loop. There’s a couple of useful terms that came up, and do get legislated in the case of lethal autonomous weapons, which are like having a “human in the loop” and then also having “meaningful human control.” And I think that they should try to do both those things in this case. I’m not sure if they’re the best places to try to set the rule, but I think that there’s a lot of incentive for the companies to keep humans in the loop when it comes to these things.
Now, could the humans on the safety team tell if the AI — GPT-8 — was misaligned and was making a misaligned GPT-9? I don’t know. It’s possible that they couldn’t catch it even if it was.
Rob Wiblin: At least not directly, probably. But they might be able to use other AI tools to inspect the outputs and to have them debate one another about whether they’re faithful. There are various things that people have suggested.
Toby Ord: Yeah, there’s a lot of things. I know when Buck was on this podcast, he listed a huge array of different possible control tools, so I wouldn’t rule out that there are ways of achieving that.
But it might be that, of the four different ways I outlined, it makes things riskier to do recursive self-improvement. It might be that having a human in the loop as one of the possible remedies doesn’t help with this particular one of these ways that things get riskier.
Rob Wiblin: Back in 2023, there was a very famous letter suggesting that we should have a six-month pause on training new, more powerful models than GPT-4. And I guess in the meantime the request was that the companies get their house in order and figure out internal practices that would allow them to more safely train and release models going forward.
I decided not to sign the letter at the time because I wasn’t sure whether it was a good idea. I think with the benefit of hindsight — I suppose hindsight is 20/20 — but I think it would have had almost no benefits, or at least the pausing on the training I don’t think would have been very useful because it wouldn’t really have stopped progress: compute would have continued advancing, they’d still have been able to do many types of experiments, and I think we would roughly have the same capabilities now as we would have given that there was no pause.
So given that, in 2023, I think it would have been a mistake, there’s a question: why is it better now? Why do it in 2026? Why not wait until 2027 or 2028? I think there’s various things that you’re trading off.
One reason that it was premature in 2023 was that the models just weren’t capable enough to actually be a direct threat in any particular way. So there was no immediate risk that was being reduced. And furthermore, during the pause, the models weren’t yet capable enough to assist you with alignment research, with governance research, with really anything. They in fact weren’t actually that useful as it turns out. So in both respects, not very much would have been accomplished.
And so in my mind, what you’re doing is trying to optimally time the point at which you have some moratorium, so that you’re actually reducing some direct risk that the AIs might pose. In the meantime, they can help you speed up a whole bunch of other work that you want to do, that you want to speed up relative to AI capabilities. But you don’t want to miss the boat completely. You don’t want to delay this so long that things go horribly wrong before you’ve had a chance to slow them down at all.
Which is a difficult question. But I feel like the case that the timing is right is stronger in 2027 than it was in 2023, and it might be stronger in 2028 than it will be in 2027. What would you make of this optimal timing issue?
Toby Ord: Yeah, I mean, it might be stronger in 2028, or maybe it’s too late.
There’s a lot of people who are framing it like this, where the idea is there’s something like a six-month pause, there’s a certain amount of pause in capacity — let’s say at six months. And then the question is this optimal timing.
I don’t really like this framing of it. I did sign the pause letter back then, and that’s partly because it had this phrase in it. I can’t remember it verbatim, but it said, as the Asilomar principles, which many leading people in AI have signed — and in fact, I was at the conference and I signed them — one of the principles was that creating artificial intelligence that surpasses humanity in general abilities would be one of the most monumental things to ever happen in human history, and should be planned for and managed with commensurate care and effort.
And I signed that, and it’s obviously true. You can’t say, no, it should be an incommensurately small amount of effort compared to how grave an issue it would be or something. And then I thought, what is the current level of effort? And the current level of effort in 2023 was not very high. The labs had alignment teams, but the teams were something like 10 people. And that is not a commensurate amount of effort when thinking about something that would be one of the most transformative effects in the entirety of 300,000 years of human history.
And so I thought, actually, yeah, that’s right. I signed this principle. The principle’s clearly right. And the principle doesn’t allow us to continue at the moment, like when we’re getting into this zone.
The letter did heavily mention six months, but it also said at least six months, and it kind of implied because it said six months to give them time to do X, Y, and Z. And it did say “at least” a couple of times. So I interpreted it as the pause isn’t ending in less than six months. But if you don’t do these things, if it doesn’t look like humanity is spending commensurate effort on this thing, then it’s going to continue until humanity does. But it was a bit ambiguous.
But that’s the kind of thing that I’d be more in favour of. I think a pause, we tend to think of it as a time-limited concept in discussions these days. So I think that a “moratorium” is perhaps a better word: it implies, A, there’s a moral gravity of the issue. A pause kind of implies we want the thing to happen, like we want to end up at that destination, we just don’t want to end up at it quite yet. Whereas I don’t think it is clear that we want to end up at the destination of superintelligence, at least not for the foreseeable future.
There’s some serious issues we’d have to sort out before we perhaps condemn all future generations that could ever live to living under perhaps the yoke of these superintelligent entities that we’re unleashing into the world. And for those conversations to spread far beyond Silicon Valley, but also beyond the West — to involve people all over the world and to hear what people have to say about it — you could never get complete consensus or complete permission to proceed with that, but wanting it to be something that at least half of the people on Earth think is a good idea, as opposed to something that substantially fewer than half the people on Earth think is a good idea at the time–
Rob Wiblin: Or have even thought about it.
Toby Ord: Or have even thought about it. It just seems kind of crazy from any kind of legitimacy perspective of ushering in this thing against the will of almost everyone.
Therefore, I think some kind of a moratorium that’s more of the format of, “We won’t go beyond some point, either until some kind of standard is met or for the foreseeable future — but people in the future are welcome to lift the moratorium if they think we’re past that point.”
Where you could just say, “Look, I’ve decided we’re not going skydiving.” But that doesn’t rule out that maybe your child then convinces you by showing you a whole lot of evidence that skydiving is actually OK. And you’re like, “OK, skydiving is back on the menu.” You’re not saying, “Forever: though the sky may fall, we will never go skydiving” or something. You can say we’re making a decision not to do it, but you could revisit that decision.
I think something more like that, rather than just, “By default when this amount of time elapses, we’re right back at it” would be the smarter approach.
Rob Wiblin: Yeah. If you talk about moratoriums with people at the companies, I think the thing that they immediately want to know is: “What exactly do you want us to do while we’re doing this pause or while we’re slowing things down, what exactly is on the to-do list so that we can then release this moratorium?”
In a sense that’s good, because they’re the people who might be having to do a bunch of this stuff, so it’s natural for them to ask, “Exactly what are you requesting?” But do you think you would want to design it with that structure, that we need to basically tick all these boxes and then we’ll feel comfortable and then we’ll go ahead? Or do we want to say we actually don’t know when we’re going to feel comfortable?
Toby Ord: Yeah, perhaps more the latter, although gesturing at some of these boxes I think would be reasonable. And there have been some attempts at this. I think Yoshua Bengio has good stuff on this, suggesting not that it would have to be absolutely impossible that the thing would go wrong, but it would be kind of normal evidentiary standards for imposing large amounts of risks on a whole lot of strangers or things like this — that it would meet the normal standards that we apply to that.
I think that there’s ways of gesturing at it, but it could well be that we don’t currently know what those set of things would be.
Here’s another one of these principles from the Asilomar principles, which I think have generally been forgotten: it said that because advanced AI systems could pose some risk of human extinction, the expected consequences of this — so like the probability multiplied by the stakes — could be colossal and things should be planned for proportionately with those consequences.
Well, if you think that there’s say a 1-in-100 risk of that — so even just 1%, so you’re 99% sure it doesn’t happen, so almost as confident as you could possibly be that it’s going to be safe — that still is 80 million lives in expectation, which is bigger than World War II, and that’s just thinking about the one generation, not even taking into account anything about future generations. So the expected consequences are bigger than World War II.
So are we putting in commensurate effort with something at the scale of a World War II? We’re getting towards the perspective of putting in commensurate effort to World War II in creating the risk: all these people going out, building these data centres as quickly as they can and investing giant amounts of money. But in terms of actually avoiding the risk, and trying to actually pay appropriate attention to all of the human lives that could be immiserated or killed in this thing, not remotely commensurate to World War II, in terms of the effort. There aren’t millions of people who are working on trying to avoid this thing. It’s bigger than it was 10 years ago, but it’s very small.
And so I signed that principle. The principle’s obviously true. You can’t say, “No, we should have an amount of effort that’s incommensurately small compared to these stakes and the people who would be threatened by our technology.” But clearly we’re not meeting it.
You know, these principles were signed by Demis Hassabis and Dario Amodei and Sam Altman and Elon Musk. They all signed them. They’re also obviously true. You can’t say, “Back then we were so naive, we didn’t realise that you shouldn’t have incommensurately small efforts compared to the stakes” or something like that.
Not that you can bind people or something with this, but I think that you can use that to lead a discussion. There was a recent call for a moratorium on superintelligence, which didn’t really clearly define what threshold that would be, but the point of it — and I signed that — the point was to say, here is a kind of vague area of AI systems that far surpass humanity across the board at all cognitive tasks. And to say, we are not remotely ready for that. Should we begin the process of working out exactly what lines to legislate on and so on, because we definitely don’t want to go there?
And the idea was that people would say we definitely don’t want to go there. Exactly how far down the path we want to go is up for debate, but let’s start that debate. Let’s form the agreement that we don’t want to go to the end of the road, and then we can work out how far along the road it would be reasonable to travel.
So I think that you need to have those conversations. If people say, “But why not one step further along the road?” or something, and they try to catch you in some little paradox or something. I feel that that is missing the point. There are people who will just walk right off the cliff.
Ways a moratorium could backfire [00:53:32]
Rob Wiblin: I do definitely worry that if we impose some sort of moratorium too early, that there will be a significant backlash to this and people will view it as a failed effort and people getting far too anxious, getting anxious far ahead of time. And eventually the moratorium would be lifted, and this would then make it more difficult to do it at the appropriate time, when the actual direct risk and the ability to speed up the research that you want is much more promising. Do you worry about that?
Toby Ord: I think it’s a concern. It could well be true. For example, one way to say what you’re saying is that the longer you run the moratorium, the harder it is to keep it going or something. And so there’s some kind of counter-reaction to it which then makes you lift it, and perhaps lift it with it blowing some steam off from the thing you were trying to suppress.
But I think another possibility is that the longer you have the moratorium, the easier it is to keep it. And I think that’s probably right.
Rob Wiblin: Because it’s normalised?
Toby Ord: Yeah, it’s normalised. And also people have stopped doing the other thing. The people who had overinvested on these extremely bullish projections about what they’d be able to do if there were no brakes, they’ve all written off their losses, and other people aren’t champing at the bit to follow them and make a whole lot of investments that turn out to not be cashable.
So I think it’s not clear. Some people say it’s this limited resource that we should be very careful when we use it and so on. And in a world where it’s also very hard to know when that optimal timing would be, it’s not like we just know the optimal time — and also once we say “This is it, this is the optimal time,” how long will it be before it actually happens? It could be quite a while. So we might need to say, “Let’s go now” so that in a year’s time they actually do it.
Yeah, so I think there’s other people who say it’s like a muscle that the more you practice, or the more you’ve had this moratorium, the better you’ll get at it. I think they’re both possible, but we don’t know. So it could backfire and we should take that seriously as a possibility. But by the same token, there’s no certainty that it would backfire.
Rob Wiblin: Another way that I worry it would backfire is while you’re under the moratorium, unless we have a bunch of other rules as well, the supply of compute continues increasing enormously. Like we’re still going to be printing lots of chips and making lots of fabs and all of that, doing further research on chip design and so on. And so at the point you later release the moratorium, you would expect a whole bunch of catch-up growth as you train bigger models or now use the prohibited techniques because many of the underlying driving factors that were causing the progress have continued underground in the meantime.
And furthermore, on top of the fact that it might go faster at the point that you release it, you would expect that if you had a moratorium on particular practices, that you might get basically a whole bunch of different actors, a whole lot of different companies converging up to the level above which you can’t go. And so it’ll be like a more multipolar, a more competitive situation at that point — which could give groups even more reason to try to rush ahead. What do you think?
Toby Ord: It could be. I don’t think that’s impossible. Once you’re worried about that, there’s ways of adjusting the thing that could help deal with it.
For example, if ASML stopped producing extreme UV lithography machines, it’s not clear that there would be a growth in the amount of chips that the fabs can be producing per year and so on. So there’s possibly just one single company in the world, who being told that they can’t do the thing that they’re doing, would actually stop that issue. People often throw up their hands, like how can you possibly deal with this? And it’s like they don’t realise quite how concentrated some of these skills are.
Rob Wiblin: It’s basically just a handful of companies. I guess it does feel like a heavy lift, at least at the moment, I suppose, to imagine this.
Toby Ord: Yeah, but again, it’s the question, even if the people at these companies start to think that they will just die if they keep producing the things they’re producing? So we’ll see how far we go down that road. But it might actually not be that heavy a lift. Or again, the people who regulate them might have to just think, you know what–
Rob Wiblin: “We’re just going to bite the bullet.”
Toby Ord: Yeah, exactly.
We should just ban unmonitorable chain of thought [00:57:46]
Rob Wiblin: I think ahead of getting any moratorium on superintelligence or RSI, I think a type of moratorium that is plausibly on the table, I think in the next year, is a moratorium on a particular dangerous research practice that I think many of the companies are not enthusiastic about, which is unmonitorable chain of thought. So eliminating the ability to inspect what the models are thinking as they’re reasoning about stuff.
There are various research avenues you could go down that would make it impossible to read what they’re saying, or at least much more difficult to read it. And I guess there’s ways of making each forward pass so long that they could actually engage in quite a substantial amount of reasoning before you saw anything that they output.
I guess many different senior people in AI at different companies have said that they think this is not a great path to go down. And as far as I know, there’s no company that has gone very far down that path. So there would be no one who’d be super disadvantaged if they all could come to the table and say, “We’re going to set that one aside for now, because this is probably the worst risk/reward of any of the options that we have available.” So it’ll be an interesting step that we could take towards beginning to limit what research practices we permit.
Toby Ord: Yeah, I think that’s a good example. I remember at the time when one of the papers had just come out talking about how important this was, and my response was that people should take this really seriously.
If someone creates some new method — and publishes it and so on, and it gets used — that makes chain of thought illegible and gets some kind of substantial advantage out of that in terms of capabilities, then it’s possible that that’s the worst piece of research that’s ever been done by any human — worse than inventing chemical weapons, or various other things that have happened. They should really take this seriously, kind of like: “Are you the villain of the history books?” If we even have history books.
This is a big deal, and people should be very careful about it. And that feeling in their bones should be more common. And I think that this is an interesting case actually, because at the labs it has started to solidify over time. I think the longer you go without someone doing it as well, the more it seems like this is a norm, right? I think you get this kind of norm solidification.
A bit similar with the nuclear taboo — where for a couple of years after the end of World War II, it was just: we’re not in a major war, that’s why there haven’t been more nuclear weapons dropped. Then when the Korean War finished without more nuclear weapons, it was like: actually this seems to be a thing that we don’t want to break.
So, yeah, I think that a norm is kind of growing around this, and it would be good if it was solidified. There’s a lot of ways of doing that. There have been various papers with coauthors from different labs and things that help to try to solidify it. You can move beyond that to try to create industry best practices or standards. There could be standards and best practices around not training on chain of thought that also help to delineate some of the related issues, such as perhaps systems that can do more in a single forward pass, also making things less monitorable.
There’s a number of different approaches. I think getting more clarity on what they all are and then building these agreements, I think that it would be great if people did more work on that — of actually bringing people around a table, really getting the people to look each other in the eye and realise just how important it is and build up that trust that, yes, we need to push on this and we need to get to some point where we agree not to do it.
Maybe it possibly could have an escape valve on it, in terms of we will not do any research on this, but if someone else is out there with systems that have unmonitored chain of thought, then we’ll do it. “We won’t be the first to do this” or something. That could be a way that they’re more happy to sign up on, or something, assuming that they can actually tell if someone else has broken it. I don’t know the details about how hard it is to verify this one. It could actually be quite challenging.
Rob Wiblin: You might require whistleblower rules, I guess.
Toby Ord: Yeah, that could be the way to do it. But I think that would be great. I do like the example. It’s a smallish one, but it’s very focused, and it really does have bad risk/reward tradeoff and so on.
Rob Wiblin: Yeah, I feel like I’d really love us to urgently codify this, because I worry that you could have some kind of rogue researcher who I guess isn’t worried about any of this and just pushes ahead. And once you know the technique is there, once it’s so easy to implement, then I guess one actor might do it and then they’re like, “Well, I’m taking a relative disadvantage if we now agree to this.”
Toby Ord: That is indeed what they’ll all be saying. It’s also an interesting example because I think that things have kind of turned on… Not from everyone, but it’s like a lot of people don’t want to train on chain of thought.
They’re not thinking, “If only we could do this and get away with it, or something, that would be great. But we won’t because other people will…” They kind of find it distasteful or something. By this amazing fluke, compared to all of the thinking about this 10, 20 years ago, the AI systems, by and large, think in English and we can kind of observe their thoughts. If you’d said that to people 10, 20 years ago, they would have just been like, “It’s insane. There’s no way that that is true.”
Rob Wiblin: It would have seemed miraculous.
Toby Ord: What’s your best verifiable, or your best interpretability—
Rob Wiblin: You just read their minds!
Toby Ord: Like, “Oh, you just look at what they’re thinking.” It’s like, really? And yet we got that, it wasn’t even a deliberate alignment technique, I think.
I think that was something that’s just an absolute gift, like manna from heaven for the people who care about the safety of this technology. And giving that up or throwing it away for — I suppose it’s like a factor of two compute saving or something. That’s probably the kind of thing people would throw it away for. But you get factors of two compute savings every couple of months, through some kind of efficiency savings that new cache rules bring or something like that.
But maybe we will be so dumb as a collective group to throw it away for something so small, but we should definitely try not to in building these things. I think that once you get these norms changing, and once you get that people don’t even want to do it, that’s maybe the key to a lot of these discussions: is that lots of people are imagining people are champing at the bit in order to do this behaviour, and they’ve got all of these incentives, they make a whole lot of money, or they avoid going out of business when their rivals are getting too far ahead of them, or something like that.
But if you can make it so they don’t want to do the behaviour, if you can win the moral argument — that it’s bad, you’re a bad person if you do this behaviour; possibly one of the worst people who’s ever lived actually, if you create this technology — that changes things.
If you ask why isn’t anyone cloning any humans, or why isn’t there a whole lot of genetic engineering going on creating genetically superior humans? If you went back and looked at what everyone was saying from 1880–1920 or something, every intellectual group was talking about eugenics and it was just generally thought that, once the capabilities were there to understand how to actually change the makeup of people in future generations, that everyone would be doing it.
And I think they would be pretty bewildered to find out that we actually gained the abilities to do that and then just no one’s doing it. And they’d say, “Is that because you’ve got all these verification techniques and you’re in the labs in all these different countries just making sure that no one’s doing it?” And it’s like, “No, we did this thing–”
Rob Wiblin: Nobody wants to.
Toby Ord: We stopped wanting to do it. It’s like, well, how did you do that? And the answer is moral change and norms change.
People really undervalue this. I think that this is a different approach to trying to use narrow incentives, to try to resist the massive corporate incentives to keep building more and more powerful systems. That’s a really hard battle to win. But actually changing the norm so they no longer want to do something can definitely work.
Rob Wiblin: Yeah, you’re so idealistic, Toby. I feel like this kind of idealism is rare in AI discourse at the moment. Or the idea that well, we could just, for moral reasons, not do stuff. I don’t know. But there’s this analytical frame that you get into where it’s like, “Well everyone is going to just follow their incentives to do that.”
Toby Ord: Yeah. I think it’s rare in the world in general to think this way, but I think it’s powerful. And I feel like if you were an economist and you’re trying to understand, say, women’s suffrage in England back when those debates were happening, you would have thought there’s this currently kind of privileged class of people — men, or if you go back further, landowning men — and why would they share their franchise with other people? Currently they’re holding almost all the power. Why would they ever do this? And I think that something along the lines of “it’s the right thing to do” is actually a key aspect of what made those things happen.
Rob Wiblin: They’ve almost gone too far.
Toby Ord: I think it kind of has. I think it is part of the moral picture, but–
Rob Wiblin: It’s a bit overweighted, perhaps. And I feel that someone who is just predicting everyone’s behaviours as like narrowly self-interested, kind of psychopathic people in these economic models, I think they would have made bad predictions. What they do is then eventually, maybe in their utility function, one of the things they care about is like their wife, or one of the things they care about is justice, or maybe justice gives them a warm glow which is one of the forms of happiness. They can have a Mars bar or they could have some justice. They have these ways of trying to incorporate it in, in order to make sense of something after the fact, but they rarely think of that in prospect when they’re trying to predict whether something will happen.
Toby Ord: But it turns out that, “Why didn’t I do some thing? Because I thought it was deeply wrong” is often the answer.
Also that we can change what we think is wrong and we can learn that something’s wrong. My favourite example is with environmentalism: that up until say the 1950s or so, it just wasn’t part of what people thought ethics included. Then there was this very big change over the 1960s and 1970s, to the point where I think in the moral education at schools in Britain, the main moral education is kind of don’t litter, carbon dioxide, global warming, a whole lot of these things.
It’s maybe a bit overweighted, but it’s hard to remember how underweighted it was. But if you read certain things, like I was reading children’s books, I think Richard Scarry from the early 1960s. Boy!
Rob Wiblin: I don’t know this one.
Toby Ord: It’s like Busytown or Happytown or whatever, where there’s all those little animals going about their day. Then on one of the pages, they’re like, “Let’s build a road to this factory.” And there’s this nice little wilderness and they bulldoze it all down and they build hotdog stands along the road, and it’s just–
Rob Wiblin: And they celebrate this?
Toby Ord: Yeah. From today’s perspective, that would appear as a parody of this evil person who’s doing that. But that was just like, “Hey, we’ve now got our road and we can have a hotdog on the way.” And then the factory is belching out all of these fumes and things.
Rob Wiblin: It’s unimaginable in a children’s book today.
Toby Ord: Exactly. And it’s just a generation ago or maybe two now or something. But you know, moral change can happen in a big way, possibly even going too far or something. Yeah, people really neglect this.
Why Toby thinks AGI is a decade away [01:09:28]
Rob Wiblin: A lot of prominent people in the AI world think that AGI is coming very soon, like in the coming years. I know that you think things could take significantly longer. Why could it take until 2040 or even longer than that to get to AGI?
Toby Ord: There’s a bunch of reasons. A lot of the bullish projections at the moment are driven by the possibility of recursive self-improvement, and it’s possible that doesn’t pan out. That’s one possibility.
Another is the possibility that we decide not to do it, so that there is some kind of global moratorium or something else like that, where even if we could, we don’t. And in that case, I think that it’s just a very credible way that you get at least a 10% chance that it doesn’t happen by 2040, to do with choices made not to do it and some kind of treaty that then has some verification or something that’s actually blocking it.
But even if neither of those… well, I guess let’s suppose that there isn’t such a moratorium. It could just be that it’s a lot further off than we think. In the world of AI, a lot of the progress is tracked through benchmarks. You start off at just solving a few percent of the problems in the benchmark set. Let’s say it’s image recognition. And it started off with this MNIST [Modified National Institute of Standards and Technology], which was identifying individual digits, handwritten digits, like 4. And then progress kind of went from getting a few percent of them right to getting almost 100% of them right.
And then they started off with ImageNet, where it was more complicated colour photographs of little objects and things, like a cup. And it had to work that out. It went from just getting a few percent of them right, and it only went up through 50%, and then got to saturation.
Then you have more challenging image recognition tasks. For each benchmark you can often predict: it looks like it’s curving upwards now and it looks like maybe in a couple of years we’ll saturate out this benchmark and it’ll be getting almost 100% right. But it’s really unclear how many more benchmarks there’ll be before you get to where you want to go.
I used to think it’s just another benchmark away or something. I felt like there won’t be that many. And now I think, could there be 10? How many times do you have to go through this process? It’s very hard to actually know.
That’s one aspect: that people who think that we’re very close to getting something, where does this knowledge come from? That if we just go through this process again, say with mathematics or something, and this AIME [American Invitational Mathematics Examination] — that benchmark was amazing for the AI community, because it took them basically through the reasoning-models era, from late 2024 through to the end of 2025, they went from a couple of percent on this thing to saturating it. So they climbed this benchmark. But then at the end of that there’s still so much more to go. So that’s one aspect.
Another one is, if you think about something like the METR time horizons: this question of how long a task — in terms of how many hours it would take a human to do it — can the AI succeed half the time if it deals with it? We’ve seen this exponential progress in this number. Then people often think, once it hits a full workday, it can do eight hours of human tasks, an eight-hour human task, then it’ll be able to just repeat that or something each day. But it doesn’t quite work like that. And then they thought, maybe 40 hours? It’s like you can do a whole workweek or something. But it’s just not clear why any particular number corresponds to reaching human level or something on this metric.
So it’s something where we’ll have this graph and there’ll be people who say they believe in the god of straight lines on graphs or something: that if you see some kind of trend, that you should be betting that you can keep extrapolating the trend. Let’s just grant all of that and let’s suppose that you can completely extrapolate this trend. Well, how high does it have to get before we reach this kind of AGI level or transformative AI or superintelligence or something? Well, no one knows. Even if you grant the heroic assumption that you can extrapolate the trend as far as you want, you’ve got the straight line, eventually if it hit this level, you could read off the time at which it gets there — but we don’t know what level it would be.
And that leads to this big uncertainty. Even when we’ve got the data, none of the data comes with a thing that then you can just read off the year from that data. That if you just want to follow the data and follow the projections and do as little second-guessing as possible and just see where the data leads, it just doesn’t lead you to an actual year when this will happen. So that’s one of the reasons to really have uncertainty about it.
Rob Wiblin: I feel like that line of argument certainly gets me out to 2030. It probably gets me out to 2035. And then I’m like, 2040, 2045… I just have a degree of disbelief that things could take that long, given the level that the AIs are operating at now. And by 2040, even on conservative projections, we would have 10,000 times as much compute, maybe much, much more. We’d have done so many more experiments, there’ll be many more people working in the industry.
Do you share? Is there any intuitive disbelief or astonishment you have at the idea that it would take until then?
Toby Ord: Not really. One aspect is that the rate at which, say, models were getting 10 times bigger or 10 times as many parameters or 10 times as much compute being needed to train them and so on, that rate — like how many years until the next 10x — has started to slow down. And there’s reason to think that there’ll be a bit of a kink in the curve at 2030 or so, when a lot of the growth has not been how fast can we scale up the compute being manufactured, but how fast can companies open their wallets?
Rob Wiblin: And suck compute away from everyone else.
Toby Ord: Yeah. Part of it is: how quickly can we convince venture capitalists to invest a whole lot of extra money, or convince people we’re borrowing from, or convince the CFO of our company to make what would seemingly be outlandish investment decisions?
Then part of it has been how quickly can TSMC [Taiwan Semiconductor Manufacturing Company] convert the share of their manufacturing that’s GPUs instead of being chips for iPhones and things like that. How quickly can they increase that share towards 100%? But they can’t just make 10 TSMCs very quickly. The time needed to have 10 copies of their entire manufacturing footprint or something is really quite large.
And so a lot of this kind of early growth does taper off. We’ll be ripping through the orders of magnitude in scaling more slowly over time. We also run into this issue that, even if you have more compute for pretraining, the amount of human data is limited and we have already gone through all of the scientific papers that have ever been published, and all the books that have ever been written up until 2025. And then if you say, “What about the books from 2026, how much is that going to move the needle on the sum total of human knowledge or something?” Maybe not very much. So data limitations and so on.
There’s a bunch of reasons to think that, conditional upon scaling not giving you some extremely powerful level of AI by 2030 or by 2035 or something, maybe that’s a suggestion that actually you need something more than scaling at that point. And I think that there may well be additional capabilities that are needed, that may require actual research breakthroughs rather than just tweaking the algorithms a little bit.
Even superintelligence needs work experience [01:17:49]
Rob Wiblin: Our new host on the show, Tom Reed, recently wrote a piece arguing that he suspects that it’s not possible to have a recursive self-improvement loop that takes you to superintelligence just occurring in a data centre because you won’t have enough training data to become highly skilled at many different kinds of tasks. Because in order to do that, you actually have to try doing the task. You have to be deployed in the real world, in factories and in offices in order to get the training data and the feedback to become highly skilled at those things. And he thought this was going to be like a super important constraint. Do you buy it?
Toby Ord: Yeah. I read this piece, and it really frustrated me at first, and then something clicked and then I really liked it. So what I was frustrated by is that it felt to me like there was a conflation between intelligence and capability. It’s true that there might be a whole bunch of capabilities you can’t gain without a whole lot of access to real-world data and attempts to try them and learn from experience and so on. Especially the things that don’t have verifiable rewards, or at least they don’t have verifiable rewards in an artificial environment. You need to go to the real world to find out. And I thought, I don’t see why you couldn’t become extremely intelligent, but you would still lack those skills.
But that’s just really a rephrasing or something of his point, or a slight rearrangement of it. But it gives a version of it that I think is really interesting, which is that maybe you could reach a kind of superintelligent system in a data centre, where it’s training and learning a bunch of things. It gets really skilled at things like mathematics, and maybe by self-improving it gets really good at a whole lot of other things.
Suppose you did. Suppose there was no obstacle to how smart it was. And so the thing that came out of this data centre at the end of this, you could think of it like an amazingly bright student who’s just finished their undergrad degree; they’re going to go out into the workforce, and the world is their oyster. They could go into politics, they could go into business, into tech, into finance, into journalism. And let’s suppose they’re so capable, they’re the most capable ‘person’ applying for the job in any of these areas, but they would be applying for an entry-level job. So the most capable person with the most potential for journalism who hasn’t yet—
Rob Wiblin: Ever written anything.
Toby Ord: Exactly, exactly. They haven’t yet reported on a hot-button issue and then got a whole lot of flak or whatever and had to work out how to navigate it, and they haven’t worked out how to trade off the risk to the reputation of the newspaper they’re writing for vs the integrity to the facts and so on.
I think what I would say is, even if you could train something with recursive self-improvement to become extraordinarily intelligent, there may just be a lot of skills and capabilities that it won’t yet possess. And you wouldn’t fault its intelligence for that. You wouldn’t say it’s less intelligent because it doesn’t possess these skills. Suppose, for example, Tim Cook is stepping down from the CEO of Apple and they’ve already got someone lined up, but suppose that they didn’t and they thought this AI is superintelligent, we could have it be the next CEO of Apple. Well, I’m not sure about that. In the same way as if there’s a really bright student who’s just finished college.
Rob Wiblin: The most fast-learning 21-year-old, I guess.
Toby Ord: Yeah. You’d still say no. Actually there’s a bunch of experience about the culture of Apple, for example, that you would need to know in order to do this job.
And also a lot of relationships and trust that some people have built up. It helps you realise that it’s connected to the diffusion question of how much of a lag is there from having these amazing AI capabilities demonstrated to them, to actually being out everywhere in the world? Where it could be that to be a great CEO of a Fortune 500 company, that it does actually take a decade or more of just calendar time in terms of approaching new opportunities, seeing the business cycle happen, seeing people who are betting the wrong way get wiped out by the market, and it’s not just cases where you can learn from what’s happened in the past. There’ll be a whole lot of new conditions that have never existed before, including with AI totally changing everything.
So it made me realise that, instead of saying you can’t become superintelligent in a data centre, I’d say even if you become superintelligent in a data centre, that doesn’t necessarily make you super capable and able to do all of these jobs. What it would probably leave you with is being a great entry-level employee in anything, and advancing along their career trajectory faster than a normal human would. So you’re better than a relatively fresh-faced 21-year-old of any stripe or something, but not that you’re better than people with 30 years of hard-won experience or context on those jobs.
That could mean that that doesn’t say much about how transformed will the world be in 50 years’ time, but it does say that actually maybe if you had this amazingly intelligent AI in a particular year, maybe a couple of years later, the world isn’t that transformed because it can’t be doing most of the jobs yet.
So that did make me think about an additional delay to add into this kind of calculation of: at which point do you get recursive self-improvement, if it’s possible? Then how long does it take to reach a of highly intelligent system? OK, but then there’s another delay between how long does it take to have the highly intelligent system, let’s say, and it getting into a whole lot of companies.
Rob Wiblin: And having lots of concrete skills.
Toby Ord: Then once it’s in those companies, how long does it take from entering that company to understanding enough about these types of roles in order to be able to really deliver at a senior level of performance? It may add, I don’t know, somewhere between six months and 20 years, depending on the particular job of those areas. I thought that was very interesting.
Rob Wiblin: Yeah, I thought it was a very interesting point. I guess Tom sounded pretty confident that this was going to lead to big delays, and I think that might be right. But I guess I also think it might be wrong for a couple of different reasons. You could have really big improvements in the speed of learning basically by that point, or the sample efficiency is the technical term for how many examples or how much experience do you need to gain to figure out how to do stuff.
I guess human sample efficiency is actually unfathomable, extraordinary in some ways. I don’t often get to pull the parent card, but having a toddler, it is unbelievable that they can see like an example of a pig and then they’ll be able to tell all of these other things that look very different are also pigs. And this picture, this stylised picture of a pig is also a pig. I don’t actually understand how on earth it’s going on. I guess AI researchers don’t know either.
Toby Ord: It is incredible. I went through the same thing with my child.
Rob Wiblin: The out-of-sample generalisation.
Toby Ord: And you know, these drawings will be in the standard way that we depict things, we don’t teach the children how we depict things. But it is these line drawings, and things look nothing like line drawings. The line drawing is this extremely stylised thing, where the edge of a mug or something, that’s the only bit that you draw and you don’t draw anything in between. But there are actually no lines there if you look at the world. Yet they’ll see a drawing of a teapot and then they’ll be able to recognise one in the real world immediately. It is very remarkable.
I think that the upper bound on what is possible to learn with Bayesian systems and various other things is even better than the humans on this. But for deep learning, it is really an Achilles heel. Before it showed that it was just so successful once you pay the costs of doing all this extra training, a lot of people were dismissive of it because of this lack of sample efficiency, and that hasn’t gotten much better.
Rob Wiblin: Yeah, so that’s a point against what I was saying. They have awful sample efficiency but you could get like a million-fold improvement. We know it could be a million-fold better because humans are a million-fold better and we’re presumably not the very best that’s possible.
That’s one way that this sort of thing could not pan out is if, through recursive self-improvement or some other way, you have a different architecture for learning.
Toby Ord: It could be that they can learn from each other. It could be that there’s a lot of complicated things about being a partner in a law firm or something like that. But each year there’s like 10,000 deployed models that are learning and pooling their information. You could have a situation like that where that’s also enabling it to learn faster. Although if the markets are correlated and so on—
Rob Wiblin: Still only have experience of that year.
Toby Ord: Yeah, that will be helpful for rare things like a particular type of client who comes in and is really obnoxious or something. And how you deal with that rare case, it’ll be helpful for that like pooling it across many different firms.
But it won’t be helpful for the rare, like once-in-a-decade or once-in-30-years, events that happen and upset the whole thing. So it is by no means certain, but it’s quite possible that there would be substantial delays, at least for them doing these non-entry-level tasks.
Rob Wiblin: Yes. So that’s one way, is that the AIs could learn very quickly because many, many instances of them are deployed and they all collect data and they pool it together.
Another one is we’re already beginning to put video cameras basically on people’s heads and record all of their keystrokes. It’s possible by this time we’ll have actually a huge new data set of basically what lawyers do and how they think and what they look at and so on that they could learn from.
Toby Ord: Yeah, that could well be right. It is interesting that Meta were kind of a trailblazer on this one and have just recently gone back on it. The employees absolutely hated it.
Rob Wiblin: Oh, really? OK.
Toby Ord: It was also clear that they were just literally automating. It was a humiliating way of being automated out of your job, with like these cameras just watching it. The level of demeaningness to the person and so on is a real kick in the teeth for these people. They’re already pretty demoralised over there with falling behind on AI and so on. So they hated it.
Then there was, lo and behold, a big leak of information because it was logging every single thing that they did. I don’t know the details of it. But then when that happened, they were just like absolute revolt on this issue and we’re getting rid of it.
If we’re in a world where there’s increasing protests on the street about unemployment and so on, being driven by this — then the answer is–
Rob Wiblin: “We’re sticking cameras on all of your heads.”
Toby Ord: All the remaining lawyers who haven’t yet been automated in the firm and now have their cameras on them and so on, because the owner is planning to replace them next year. You know, people may not wear it, is what I’m saying. It’s possible. I don’t know which way that would go because there could be large financial incentives for the few firms who do it. It could be possible. Maybe you pay them larger than their entire salary in order to wear this camera because of the amount you’re making.
Rob Wiblin: It’s like a redundancy payment.
Toby Ord: Yeah. But if the public sentiment is just put sand in the machines–
Rob Wiblin: I could imagine it being banned.
Toby Ord: Yeah, it could be banned. It’s a good point. It might be that companies that try to do it are just shamed. I don’t know. But by the same token, I think there’ll be some of it.
Rob Wiblin: It’s another ‘it might or it might not.’
Toby Ord: Yeah, it’s hard to predict.
Rob Wiblin: And just to finish out my various objections, I think the last one was: some things I imagine you can learn through simulation, and I guess they already kind of do this. I guess that RLVR [reinforcement learning with verifiable rewards] is a sort of simulation. I don’t know.
But yeah, you can imagine if you have a really good internal world model, then you might not have to go and actually do these things in the real world because you can just imagine it. I guess it’s like I think humans learn partly through dreaming, and I imagine there’s some of this is like we imagine situations and react to them and that helps us to process the information that we have. And I guess perhaps AIs could do that as well.
Toby Ord: That’s right. Although there’s also another interesting case where, when people are trying to estimate how much they’re being sped up by AI, they tend to think of their hours in the office, so on-the-clock hours. But if you get an AI to do some task for you and you save four hours and it does it instead, you don’t learn from having done the task. So you’re not necessarily in as good a position as you would have been if you’d done it.
And then also there’s a bunch of thinking, certainly in academia, that people classify as like ‘shower thoughts’ or something — where you’re still only having one shower a day, even at Anthropic, where they’re being sped up by 4x — and there’s just certain kinds of processing of what’s going on. And perhaps more realistic than showers, although that’s perhaps where some of the thoughts come out because you can’t be reading a book.
Rob Wiblin: This is the one time you don’t have a screen in front of you. For now.
Toby Ord: Or similarly, going on a walk or something is another kind of famous example. But sleeping, you know, you still only sleep for eight hours. They’re not getting like three nights’ worth of processing of everything that you’ve been tracking and having it click into place in your mind or something. But a lot of — if you read biographies of scientists who’ve had major breakthroughs — it seems to be some of these aspects of they finally had a change in scene, they went on holiday or something, or they woke up and the thing was just clear to them, that maybe it had been processed.
We know that the brain does an awful lot of stuff during these dreaming processes and so on. And presumably it’s not nonsense and it is actually part of what makes this work.
Rob Wiblin: If you deny people dreaming, I think their ability to remember what happened the previous day is catastrophically ruined. So it’s a huge factor in learning, I think.
Toby Ord: Yeah. It’s worth noting that we’re not speeding those things up, and so it might be that some of the numbers that we’re getting are just really not the relevant ones.
Is AI coming for mathematicians? [01:32:22]
Rob Wiblin: I’ve got this piece that we published recently where I try to make sense of all of the different updates that we’ve gotten about timelines to AGI through 2026. And one of the ones I had on the list was that OpenAI, one of their models — I think they haven’t named which one — made some progress on disproving a commonly believed conjecture about the unit distance problem, which I guess is like a possible sign that it’s becoming a creative researcher. It’s able to have original thoughts.
But I ended up concluding that I’m really unsure whether to be impressed by this or not. It’s cool, but it’s very hard to interpret. Do you have a take on whether this is impressive or not?
Toby Ord: Yeah, I think it’s impressive and I think that the mathematical community were clearly impressed. That said, it may have already been baked into your assumptions. A lot of people are assuming radically impressive progress every year, so seeing something that is really impressive may just be only enough to keep it on trend.
But this unit distance conjecture is, I think, a genuinely natural and interesting question. Ultimately, the idea is if you have a 2D plane and you can put some points on it, if you could put N points anywhere on this plane, how can you arrange them to make as many as possible exactly, let’s say, one centimetre away from each other? So putting them in a circle that’s evenly spaced would be a way to do it.
And that would let you get, for N points, you could get N unit distances, but it turns out you can get a little bit more if you put them in a square grid. And then the AI, it had been conjectured by Erdős, who’s like a very famous mathematician, that the square grid was basically as good as you could do, but then it showed that you could do something a little bit better. And so it didn’t prove a theorem, but it disproved his famous conjecture that a lot of people thought was true.
And this is unlike a lot of other results that AI systems had produced before this. This was one that genuinely a bunch of mathematicians had spent some time trying to solve. It wasn’t just like something that wasn’t solved because almost no one had spent any time on it.
So I think, pretty cool disproof of a conjecture. Unfortunately, the conjecture was quite elegant. And instead now we’re in this world where it’s like, oh, it’s more complex than that in some kind of slightly ugly way. But it did it, if you look under the hood, by combining this area of combinatorial geometry with another area of mathematics that was previously thought to be unconnected.
So in some level, it’s a bit of a shallow result because if you happen to know about both those things, like if you found out that the person who had done it was a human and they’d just been brought up in this other area where they became intimately familiar with this kind of weird mathematical structure instead of a square grid, this other kind of structure that you use, and then they were told about the unit distance conjecture and they thought, “Why don’t I use the thing I did in my PhD thesis?” And you’d be like, “I guess lucky that you had that combination.”
What was exciting about it for mathematicians was that there wasn’t a previously known connection between these two areas. It found a kind of connection which will let mathematicians import a whole lot of ideas from one area into the other area. They’re normally excited about results like that, when humans do them, and so they’re also excited here.
So I think pretty cool stuff. But there’s a bit of a feeling in the maths community that — and people vary on this — but there’s a bit of a strand of feeling that, “Whoa, there might not be any jobs for mathematicians, this really could be coming for us.”
And maybe it is, but I think it’s illuminating to think about the limits on this thing. This isn’t even a theorem; it just disproved a particular thing. And one of the reasons it hadn’t been disproved before is everyone thought it was true, so not many people had tried to disprove it. But let’s suppose that we kept going in this direction, and much further, and we had a system maybe in a year or two or something.
Let’s suppose we have a system that — if you give it a formal statement of mathematics, or make some mathematical claim — that it could prove or disprove it in a microsecond. OK, so we’re not going to get that. In fact, Gödel’s incompleteness theorem says you can’t actually quite get that. But let’s suppose you got it somehow, impossibly. Would maths be over if you can just take any claim and check whether it’s a theorem instantly? It turns out it’s very much not over. I think mathematicians are underselling the other aspects of mathematics.
Rob Wiblin: Yeah, what are those?
Toby Ord: One of them is like, why were we interested in the unit distance conjecture? It’s a pretty simple kind of claim.
But suppose you could prove one of these things per microsecond, if you started doing that now — I did some calculations on this — the stars would have burnt out by the time you even came up to the unit distance conjecture to even try to prove it. There’s just too many mathematical statements. There’s all kinds of statements like 1+1=2, 1+2=3, and so on. There’s infinitely many true mathematical statements. So you have to kind of greatly prune this space of statements to the interesting statements.
One way to rephrase that is like, what questions should we be asking? What are the interesting mathematical questions?
And it’s not clear that it can do that, that it can work out which questions to ask. At that point you’d still need humans to say, “This one is one of the ones I want to actually have the system spend its time on, not this exponentially growing thicket of uninteresting questions.”
So that’s the first step. Then beyond that there’s even richer things, there’s this question of: the AI systems can take some formal statement in some theory of mathematics — in this case discrete geometry — and then try to prove it, but they can’t create new theories of mathematics. So if you go back in time, say before Claude Shannon invented information theory: now, today, we could ask questions about the information that can be carried over a noisy channel and optimal coding and stuff like that. But we didn’t even know how to ask those questions back then. We didn’t know what kind of formal statements to have that would correspond to these informal and inchoate ideas that we hadn’t fully pinned down.
So when mathematicians invent areas like that, there’s a kind of extreme creativity. If you go back to the 19th century, there’d been thousands of years of geometry, and then in the 19th century mathematicians worked out that you could ask questions about things beyond three dimensions, like a four-dimensional cube: how many corners would a four-dimensional cube have? But they didn’t ask these questions prior to then. They also worked out, in the 19th century, you could have curved spaces, and they didn’t realise that we actually live in one, but they just thought it was an interesting mathematical question. They also kind of worked out about fractional dimensions, like things between two and three dimensions.
And there was this burst of creativity that you could ask all of these types of things. Or if you think about the origin of calculus, that we could ask about not just how high up is some curve, but we could ask about the slope of the curve and how that changes over time and so on. And that once they came up with these theories, all of a sudden there was this whole infinite range of new interesting questions you could ask. And the AI systems haven’t shown that they can do that at all.
So I think it’s instructive to see that, even if mathematicians spend a lot of time proving things, there’s these other layers that haven’t really begun to be automated.
Rob Wiblin: So they can’t do it now. But I would expect that stuff is coming soon. Do you agree?
Toby Ord: It might be, yeah. I’m not claiming that it won’t be able to do it, just that there’s this kind of thing that we think we can see this trend. We’re laser focused on this issue about proving things or something. Then we think that a mathematician, that’s what they do, and that it’s going to automate that away. But it’s easy to lose track of the fact that there are these higher-level questions, which I’ve always thought are the more important questions in mathematics. I’m less impressed if someone proves a difficult theorem than if they invent a new branch of mathematics where entirely new questions come into view and new concepts. That’s always been what I thought was the more impressive thing.
If we look at the history of automation of mathematics, for most of the time — until the 20th century actually — mathematicians spent something like half their time doing numerical calculations. Then the calculator automated all of that away, so it automated half of a mathematician’s job. And we don’t think that was a great shame or something. We also didn’t think we were on the verge of a singularity or something when we did that.
And then from about 1980 to now, symbolic manipulation, like solving algebraic equations and integrals and things like that, has also got basically entirely automated before the AI era — and again, that was then what mathematicians spent a lot of time doing. Now they don’t have to do it at all. They don’t even really talk about that very much.
But I think that, again, they were by and large freed from a relatively pedestrian part of their job — and that maybe if proving things gets automated, they’ll also be freed to be asking the questions and inventing these new theories. And these are areas where creativity is needed.
I think that those lessons aren’t just relevant for the mathematicians listening to this, but are potentially relevant in a whole lot of jobs. Famously in the case of radiology, the reading the scans bit we tend to, from a distance, think that’s what a radiologist does, but it turns out that less than half the time of a radiologist is spent interpreting scans. And it’s like, “Oh.”
I think there’s similar things with a lot of jobs, where we look at some component of it and see this curve of it being automated away. But we forget that there are richer and higher-level things, where it’s not that the humans are necessarily being pushed into the gaps of like… For example, some mathematicians talk about maybe humans will still be needed to explain to other humans what the new AI mathematicians have achieved. I feel like that is a case of being pushed… I think actually science communication is really good–
Rob Wiblin: You wouldn’t feel central to the story, I guess.
Toby Ord: Yeah, they would feel less happy about that. Whereas if they’re instead inventing entire new branches of mathematics, then the AIs are helping them explore them, I think they would actually feel pretty good about that.
We don’t know how long it will take before it can do those things. It could be that it’s just another year after or something like that. But I’m more pointing out that we don’t know, and that it’s a bit like with these benchmarks: we’re seeing proving start to get off the ground, where it’s starting to be able to prove nontrivial theorems. Maybe that will saturate and it’ll be able to prove all kinds of challenging theorems that would take humans years to do.
And yet we’d find out there’s another benchmark above it, which is asking the right questions, and then that starts to saturate. And there’s another one, which is inventing new fields. I think we’ve seen that a lot in the history of AI, that we start off thinking something like chess is… “AI complete” was this kind of term, which means it’s as hard as anything, such that if it can do that, it’ll be able to do everything because it’s like the last rung to fall or something.
Then people go and work on it and then they find out, oh no, actually they still can’t understand English at the time when they can master chess. Then we get to the case of Go or something, and it turns out they still couldn’t speak English fluently at that point. Then we think maybe speaking English is the thing. And then it’s like, oh no, we’ve got systems that can quite fluently deploy language, but they’re unable to do some other things.
So we genuinely don’t know how far that process goes. And there are very few attempts to systematically say, “Here’s the set of things and here’s how there isn’t much left.” Certainly I’ve found, as someone who’s generally been — compared to the average, actually pretty bullish about AI timelines — that I’ve noticed–
Rob Wiblin: How many times there are more things to go.
Toby Ord: Exactly.
The case for broad timelines [01:45:00]
Rob Wiblin: You wrote this piece not too long ago arguing that we shouldn’t think of ourselves as having short timelines or long timelines or medium timelines, but rather having broad timelines: we should embrace the reality that we don’t know when AGI is going to come.
Yeah, make the case for that. I guess you’ve already somewhat made the case for that, but is there much to add?
Toby Ord: Yeah. There’s kind of two things.
One of them is that there’s a lot of expert disagreement on this, like lasting expert disagreement. The experts just aren’t intractable. If you look overall, Helen Toner has this great piece showing that actually their timelines have shortened over the last five years quite significantly, I think by about a factor of 10 or something — at least from some forecasting platforms, have gone from 50 years to five years.
So it can change, but there’s still a very large amount of disagreement between different people. And those people come often from different fields, incorporating a lot of different forms of expertise. So obviously people who are AI specialists have very clear and relevant expertise in this, but also people who are, say, psychologists have a lot of relevant expertise. Economists have a lot of relevant expertise, if the question is will this be able to replace humans in labour and their role in the economy or something — they’ve seen a lot of waves of automation and are kind of experts perhaps at automation.
There’s a lot of different forms of expertise being brought to bear. The individuals who form their opinion are not aware of a lot of the hard-won insights that other people are bringing to the table. They’re not in a position to actually say, “I do not need to listen to your view on this.” There’s this question of what’s the chance the other person knows more than you? Half the time, if you take two experts, the other one had a better idea of the picture than you did, better or equal, but we’re assuming they disagree. And yet in order to have a narrow range of timelines, you really have to be excluding a lot of people and saying, “Nope, you’re wrong and I’m not interested really in hedging a bit by moving towards what you’re saying.” That’s one point: lasting expert disagreement.
Rob Wiblin: And yeah, actually I can’t remember: what’s the other key point?
Toby Ord: The other thing is that if you look at forecasters on this, the best forecasters I think don’t just give a point estimate for when it’s going to happen, but they explain their own uncertainty. The ideal way to do this is by having a probability density function, like a bell curve or something, although it doesn’t have to be symmetrical, that shows exactly how much probability in which year. And the lay version of this is a confidence interval — to say “somewhere between this year and this other year.” But the probability distribution is even nicer if you can get it. A lot of people have actually sketched these out. For example, the AI Futures Project — who gave us the AI 2027 scenario — they’ve got some very nice and interesting modelling of recursive-self-improvement-driven AI progress.
And the two leads on that, Daniel Kokotajlo and Eli Lifland, have got their own probability distributions you can look at. And Kokotajlo is generally thought of as—
Rob Wiblin: About as bullish as you can get.
Toby Ord: Yeah, exactly. Real short-timelines person. But if you look at his distribution — and he’s recently changed it to actually be a bit more bullish in the last few months — but his distribution, he thinks that there’s a 10% chance it will happen within nine months. That’s pretty bullish. Then his median estimate, so the 50/50 point, his over or under number, I think it’s four years at the moment. Then he thinks that there’s a 90% chance it will happen within 14 years, which is taking us up to 2040. But he would say there’s a 10% chance it goes beyond 2040.
So I like to think about these estimates in terms of what I call the 80% confidence interval. We’re often familiar with a 95% confidence interval in science, like ranging from the 2.5 centile up to the 97.5. Here I’m just thinking actually from the 10% to the 90%. So the kind of confidence interval where you think there’s a 10% chance it’s earlier than anything inside this window and there’s a 10% chance it’s later than anything. So there’s only an 80% chance that it actually happens in the window that you specify. We’re not asking for giant amounts of confidence. It’s very possible that it will fall outside this. So Daniel’s one is from nine months to 14 years. That’s pretty wide. It’s maybe at the narrower end of what I’d call broad timelines, but I think that that counts as an example.
Ultimately that’s about a factor of 20. He’s kind of saying there’s a factor of 20 uncertainty in when this thing will happen, between nine months and 14 years. And he’s saying that there’s a 20% chance it doesn’t even lie in that giant window that he’s created. So that’s pretty broad already.
If you look at other people’s, they’re also really broad. There was a summary that Epoch did of I think 10 different forecast methods, and all of their 80% confidence intervals were more than 50 years long. And I think almost all of them had more than a factor of 10 between the early number and the late number.
Then also if you look at this big Katja Grace et al. paper, where they surveyed a whole lot of experts from people who were presenting at the major machine-learning conferences that year, and they surveyed more than 2,000 people, they had another vast range where the 10% number, I think, in that case was just four years, and the 90% number was beyond the span that was being looked at — so more than 100 years. Again that’s at least a 25-times multiplier.
My own 80% interval is something like, I can’t remember, I did work this out recently, it says something like two years to more than 100 years. I think three to 100 was my estimate. So 30-times multiplier.
So people are trying to say, with these attempts to actually express their uncertainty, that there’s these wild levels of uncertainty in these things.
Rob Wiblin: In the piece you lean reasonably heavily on the fact that I guess it seems like even the most bullish-timelines people do have reasonably wide confidence intervals. I’m a little bit nervous about over-relying on that, because I guess in the case of the AI Futures people, I think they’ve spent a lot more time modelling the rapid timelines than they’ve modelled the longer timelines. So I worry that under the hood what you would find is that they’ve just said, “…and there’s a chance that it takes a lot longer and we don’t know” — it might not be as considered as the rapid recursive self-improvement scenarios that they write about.
Another reason is, maybe they’ll say it could take much longer than 2040 or 2050. I’m not sure whether they’re saying it won’t be technically feasible until then because it just requires an insane amount of compute, or we don’t have the data — or they’re saying there could be a war between the US and China that destroys civilisation. Maybe that’s why it takes 100 years. It’s very different implications, I think, between these two.
Toby Ord: I don’t think they’re saying much about the latter. In general, I worry that a lot of people who are forecasting this are largely saying, in business as usual or something, when would we have the capacity to create these advanced AI systems if we wanted to? Or something. Such that if their moratorium happens, and then we reach 2040 and it hasn’t happened — as in we haven’t reached this level of transformative AI — I worry that a lot of people would say, “Yeah, well, I’m not wrong because this thing got in the way.”
Rob Wiblin: “Because we could have.”
Toby Ord: “We could have done it.” Yeah. That’s not the forecast I think that they should be making, or at least if they are, they need to be super clear about it. Because I think the forecast that matters for most people is when will we have these systems? Not when would we, if we’d behaved optimally in a certain sense and ignored how dangerous it was, when would we have them? But rather something more like when will we have them? Although it does depend on the use.
Rob Wiblin: Yeah, maybe I’m focused on the technical feasibility question a bit more because I’m thinking this creates a deadline by which we have to have figured out a bunch of stuff, like being able to coordinate to decide whether we’re going to go ahead with it or not, or figuring out some way to do it safely.
Toby Ord: Well, it depends on the question. Suppose the question is alignment, and what you’re thinking is that should I invest in just kind of tweaking the current alignment paradigm, or working within this paradigm, or should I invest in creating new ideas: like exploring and finding fundamentally different ways of doing alignment that could have better guarantees? Like Yoshua Bengio is doing with his project.
If you’re thinking about that, it is relevant if, suppose, that there isn’t this kind of transformative level of AI by 2035 because we decided not to do it. That does mean that you have more time to actually do things.
Rob Wiblin: Figure things out and become relevant.
Toby Ord: Yeah. If instead you said, “But because we could have done it, I’m going to then ignore the possibility of these higher guarantee type methods, this exploring the space of ideas, and just focus on tweaking the current things.” I think that would mislead you. I think that for many purposes, actually what you want to care about is when will these systems exist? Rather than when could we have had them exist?
There’s also a bit of a thing that annoys me, where they’ll never know that they’re wrong if they do that version because they’re saying we could have if we hadn’t had that pause or, if the companies had just leant more into recursive self-improvement and they hadn’t chickened out on it, then we would have had it or something.
We’ll never know if that’s right or not. A lot of the idea that the forecasting community prides itself on is actually falsifiability by some date. All of the versions that say, “If a whole bunch of things that are not even fully well-specified happen, if we’re in one of the normal cases, it will happen by this point.” And then they’ll say, “I guess we weren’t in one of the normal cases, so my forecast is moot.”
Rob Wiblin: Yeah, I guess on the falsifiability point, there’s been a lot of slippage, or there’s been a lot of vagueness about exactly what people are describing. I think this is as true of me as anyone. I think I used to talk a lot about timelines to AGI. These days I often think and talk more about timelines to recursive self-improvement.
I think other people, including me in the past, we’ve also talked about timelines to artificial superintelligence. And there’s other points along that as well. I suppose because we’re all throwing out dates and sometimes not specifying exactly what we’re talking about, it’s easy to agree, but feel like you’re disagreeing.
Toby Ord: This is a big issue. Also it can explain differences between any one person’s forecast between different times. The AI Futures Project are great on this issue because they have, I think, four different versions of the question that they’re asking and they provide forecasts for all of them, and then you can see how they’re moving over time and so on. I love it, even if I think that they’re maybe a bit bullish.
But you mentioned a few there. I also think AGI is not quite the right one to be forecasting now that we’re sufficiently close to it that I don’t think it’s good to say the current models are AGI. But if someone disagreed, I can’t prove they’re obviously wrong. And it’s not unreasonable to say that actually we’ve hit the threshold of AGI, even though it can’t, for example, run a vending machine, or can’t run a cafe profitably or something. I feel like that probably shouldn’t count.
Rob Wiblin: I think 10 years ago we would have said that AGI would be able to run a cafe.
Toby Ord: Yeah, exactly. I don’t know that that one’s goalpost shifting. They might say it can prove novel mathematical results at the expert level, so you are shifting the goalpost. But I’m like, “Yeah, but we thought that by the time it could do that, it would be able to run some kind of cafe or something.”
Rob Wiblin: Thought these things would come together, and it turns out they don’t.
Toby Ord: Exactly. We partly discovered more and more jaggedness and so on of these systems that we’d thought were quite general. Part of the reason for that is — this is actually a big point — is that this process of going to the reasoning models, what that did was the verifiable tasks such as mathematics or programming, they got a huge boost because you could run this reinforcement learning loop where what you do is you get it to try the thing again and again and again and you check if it got it right. And if it did, then you kind of increase the likelihood of producing answers like that. But for tasks, such as almost everything, that are not verifiable rewards — say every other subject in the university or something like that, or running a cafe or something — it’s just much harder.
In the case of the cafe, you do eventually get verifiable rewards in the real world, but if you look at how many runs you have to do, you have to run 100,000 unprofitable cafes before you’ve actually gone through that process and spent a huge amount of money.
There’s these two different classes of skills or something. Broadly speaking, there’s the verifiable rewards and the others. And the others, there’s not that much of a theory of the case as to how they’re actually going to reach superhuman levels, whereas the verifiable reward ones, we can see how that would work.
Rob Wiblin: Is the main implication of broad timelines that people shouldn’t only make bets that pay off really quickly basically?
I guess it sounds like from your piece, your main worry is that there might be people out there who are thinking, “I really want to work on making AGI go better, but this project that I’m contemplating won’t pay off for five years. That’s just too long. There’s no point even doing it. We’ll be dead by then — or things will be out of our hands by then. I’m going to do this thing that pays off really quickly” and that thing otherwise could just be less impactful.
Toby Ord: Yeah, that’s a part of it. What particularly frustrated me from some people in the short timelines camp was this idea that they’d make these declarations that everybody should only be doing work that could have demonstrable impacts on making the intelligence explosion go better or something within a year, within 12 months, or some kind of claim like that.
That frustrated me because I don’t think it was meant to apply to everyone. I think that they just weren’t thinking about the full space. Suppose there’s someone who’s currently working in some other profession, but they’re thinking of a career change to, say, switch to working on AI policy, but it’s going to take them a year to skill up in that before they can really get started, I think that they may well want to make that career change.
Whereas if you have one of these hard-and-fast rules, it could be that if you’re already working on AI safety, that you should be primarily focused on things that can have concrete implications soon. But you want to be careful how many other people’s careers you’re somehow weighing in on when you really haven’t thought about the full space of people’s career choices.
If you look at, in my life, I’ve done things like founding Giving What We Can, where I worked out that over my life I’d be able to donate a lot of money and it would be able to save many people’s lives or create other big benefits for people in poor countries.
But by founding an organisation where 10,000 people have joined, this is something where there’s this huge multiplier that can happen. And that was over about 17 years. We may not have 17 years, but even if you look at where it was from when I first thought about this to where it was five years later, all of a sudden there were like dozens of people working together on these projects and having this big multiplier.
I think this is often the case. A lot of successful things, like, say, AI governance was just kind of like a dream 10 years ago. When I first met Allan Dafoe and he was talking about it, I was like, “What do you mean by AI governance?” I distinctly remember having to ask him, like I didn’t understand what he was talking about. And now it’s like this–
Rob Wiblin: Significant discipline.
Toby Ord: Exactly — that exists at many different universities. The people who were just getting started in it are now in very high demand from governments across the world for advice on these questions.
It’s possible to get these really huge multipliers over periods of, say, five to 10 years, such that if you say what’s the chance that everything’s moot, that AI has arrived, it’s had transformative impacts on the world of the kind where you’d want to have all of your impact before this happens. What’s the chance that that’s happened, let’s say within the next five years? I think the chance is something like 20% over the next five years. And this is a pretty high bar for transformative AI. Not just that there is something that could do what a human could do but the world hasn’t yet changed. If so, that’s like a 20% haircut on the expected value of things that you could be doing — projects that really start to kick into gear after five years — because there’s a 20% chance that’s moot. But the 20% haircut may just not be that significant if the other plan was going to have a 10-times multiplier on what could be achieved, because you’ve grown the set of people who are interested in some topic to some much larger size.
It’s just not that uncommon to have a longer timeline or project that you’re considering, or career which would have 10 times the impact of the opportunities you have in the short run, such that it survives the haircut from the fact that it could be preempted, especially with career change things.
So partly what I’m saying is that people should be very careful before they issue blanket advice to all people based on the possibility that things could happen soon. That’s what had frustrated me from the short timelines.
Rob Wiblin: I guess I haven’t heard anyone say anything as extreme as you shouldn’t do anything that doesn’t pay off for more than a year. But I could imagine that part of the motivation is just an absolute, like people are kicking and screaming because they’re just so exasperated that they feel like the rest of society — and even governments that are quite AGI-pilled — there’s no hustle, there’s no sense of urgency to address this stuff that they feel is coming so soon. Yeah, I guess it’s just difficult to balance these different audiences.
Toby Ord: Yeah, I would agree with that feeling. In fact, I think that’s kind of how I translate some of the statements that people are making, where they say no one should do this thing. Maybe they’re not imagining that everyone will start to obey that dictate. They’re instead thinking, “Too many people are doing things that take too long to pay off. And if I just tweak that slightly, that’s probably a good thing.” Yeah, that might be true.
One way I break this down is that I think that if you ignored questions about AI timelines and you looked at whether the best career path or the best kind of project to work on should pay off in the 2030s, or pay off before that, that maybe something like half of the plans that would pay off in the 2030s, once you adjust for this, you should actually switch to the shorter plan. I do think that there’s a substantial effect or something. But by no means is it saying, “Absolutely all of that is off the table and no one should write any books, no one should start any movements because all of those things take longer to pay off.” I feel like that would be a mistake.
But the mistakes that the other people are making in the other direction are much bigger.
Rob Wiblin: Bigger than that? Bigger than only focusing on one year?
Toby Ord: Maybe the issue is that people aren’t actually following that particular command, but I think that in government they’re generally paying substantially too little attention to the possibility that this stuff could happen very soon. In many other places, pretty much everywhere, their distribution doesn’t include enough of the short timelines bit.
So while I’m often focused on my colleagues and friends and people who I think have gone a little bit too overboard on overconfidence on short timelines, as opposed to just saying, “You know what, it could well be short. We need to hedge against it” for going a bit too far on it. But I think that the bigger mistake that’s been made is in the other direction.
Rob Wiblin: I think something that’s even crazier is that there’ll be governments or people in government who basically do have shorter timelines or broad timelines or whatever. But then it feels like it has no effect almost on what is going on. I guess it’s very hard to move institutions and to get them to do anything very quickly, to focus on the fact that the future could be radically different — because I guess they have a really strong immune reaction to that idea, because you don’t want them to turn on a dime.
What advice do you give to people in government about how they should approach this?
Toby Ord: Yeah, I think there’s a couple of big mistakes that they could make, and I want to be pretty clear on this. So by saying the broad timelines, I think another way to say it is: the best single summary of when will AI happen is not giving a number, like a year, but is “We don’t know” or something like that. And to embrace and acknowledge the big uncertainty of this issue compared to many other issues in sciences, where they know to within a year when it’s going to happen for some particular event. So for us, it’s not unreasonable to say we don’t know. We don’t know if it’ll happen next year; we don’t know if it’ll happen four presidential terms from now. That’s our level of uncertainty.
I think it’s good that we acknowledge it and so on, but if you had a minister for AI who’s hearing this, one thing they might think is, “You’re saying you don’t know. And that gives me permission to just assume whatever I want to assume, as long as it was in the range of things you found credible.” I think that’s often how politicians act, and many people. If you’ve got some view on some issue, like is it dangerous to eat this food that you enjoy eating, and then you find out that there’s a kind of disagreement among the experts and that some of them think it’s credible that it’s safe to eat it and so on. You might be like, “OK, so I can just carry on doing what I was doing.”
I think that’s a big mistake. It’s not giving permission to do whatever it is that you want. Instead, it’s more like you’re obligated to actually pay attention to all of the different timeframes that experts find credible. An example would be, suppose that there’s a volcano near the town that you’re in, like you’re in Italy or something, and the experts disagree on whether they think actually there’s a serious risk of the volcano erupting. Some of them say it could be within a year, and some of them say actually 10 years or more. What do you do? Well, you don’t say, “Because some of them said it could be 10 years, I’m just going to go with that.” Instead you need to be planning for both these contingencies.
Maybe you need to be saying, “OK, if it’s within one year, that’s too short a time in order to actually build defences like to divert the lava flow, so what we need is an evacuation plan for if we see signs of an eruption — how do we get all of our citizens out of town?” But then also you don’t want to say, “That’s the only thing we’re going to do, and then if it takes longer, we’ll just waste the time. We could have been building these earthworks to defend the town.”
Rob Wiblin: To redivert the lava or whatever.
Toby Ord: Yeah, there’s no particular single number as to when it would happen that you should act as if it’s going to happen at that year. Instead you should be taking seriously the possibilities that the experts are pointing out. One way I talk about this is we’re in this race against timelines. We’ve only got so long before we have to sort a lot of things out. But we don’t know if that race is a sprint or a marathon, and that makes it really challenging to work out.
Rob Wiblin: That’s unfortunate.
Toby Ord: Yeah, it is unfortunate, but we should acknowledge that unfortunate fact. And if we say it could be a sprint so we should all start sprinting now, that may not be the best thing to do.
I think overall what you find is that, in general, we should be hedging towards this possibility that it could come soon, at a time when we’re least prepared. I think that overall there is a push in that direction, and it’s great that a lot of people have been taking that seriously. But we don’t want to go beyond hedging about that possibility and instead take a full-on unhedged bet that it’s going to come soon and neglect these other possibilities, where we could have been building communities, movements. We could have been taking actions, like for example, to get an international treaty against superintelligence that’s not going to pay off in the first year or something. It’s going to take a while before you could build up that awareness.
Maybe it will take a while before AI capabilities are so strong that people feel in their bones that this is a real threat. But it could be that it’s one of the best things that we could do about AI and saving humanity from this threat. So we want to have this range or this portfolio of different approaches.
Rob Wiblin: In my timelines piece recently, I spent quite a bit of time trying to — I guess there’s so much short timelines stuff in the water at the moment, and especially among the kind of people who I imagine might listen to 40 minutes about this topic — that I wanted to do a bunch of deflating, I guess getting people to think it could take longer.
But the flip side, I think, of almost everything that I say is that it also could be quite short. There’s actually quite a number of different pathways to which we could get extraordinary capability jumps relatively soon. A couple of them are:
- It could turn out that recursive self-improvement works even with AIs that don’t have a super broad range of skills. Maybe you only need to train them on, I guess, coding, of course, but also you get them to have some decent research intuition about just AI in particular, and then they start having great ideas and insights that you can train on. Or they’re just very good at setting up experiments and basically it’s all just a brute-force search through different experiments and that pans out.
- Also, there’s this general phenomenon that the AIs are much better at well-specified, simple, structured tasks than they are at messy ones. And it’s possible that they’ll remain bad at the messy tasks for quite a long time to come because it’s harder to train. But it’s also possible that by making them superhuman at the well-structured, simple tasks, you’ll get cross-generalisation into the messy tasks. Then they’ll become roughly human level at that, and then even their worst areas are human level, and then you can go a long way from that. So that’s another possible path.
- I guess there’s a bunch of areas where the AIs are weak — like long-term planning, things that are hard to do reinforcement learning from verifiable rewards. But it’s possible that just through sheer force of effort, the AI companies could hire lots of people to basically read the outputs, read the work that the AIs do, to grade where they’ve been doing well and where they’ve been doing poorly, and basically say whether they did a good job or not. That would be expensive, but the companies are spending a lot of money. I’ve seen some people run numbers that maybe this would only double the cost of a training run on this kind of thing. So that way you could produce just enough data to basically get them to be human level at these kinds of strategic planning tasks or long-time-horizon tasks that so far we’ve been struggling to crack.
There’s probably a couple of others that I’m forgetting, but I guess I’m very worried about short timelines, like other people are, because there are a number of different pathways to get there.
Toby Ord: Yeah, I agree with everything you said. I’m also very worried about short timelines. Thinking that it may well be long timelines doesn’t make it all that much less worrying that it also may be short timelines. I think that you’re right that there’s a bunch of ways that it could come soon, especially if recursive self-improvement happens and if it’s at the easy end of the spectrum.
We don’t know whether the field of AI, in order to get to these really advanced levels, if it’s mainly just hill climbing and making small tweaks to the current basic structure and then just seeing what makes it go better and just following that gradient. Could be, and if so, then there could be really explosive growth in AI capabilities. We can’t rule out, say, Daniel Kokotajlo’s nine months. I don’t think I can rule out that it could happen in nine months. In fact, I can’t really rule out that it’s already happened behind closed doors and I haven’t heard about it yet, but I don’t think that’s likely.
What I try to say to myself with some of these things is it could well be that we have, say, transformative AI before the end of the current presidential administration in America, but we probably won’t. And so it’s useful. It’s important to know that it could happen. It could be the current political arrangements are the arrangements under which it occurs. But I think that it’s more likely that they’re not, in which case things could be really quite different.
So that’s an example of how you can learn from this perspective that you need to be able to hedge against these early possibilities that happen when we’re least prepared, but not overcommit on them or something.
I worry if companies sacrifice their principles in order to appeal to whoever’s currently in charge, or if people in the broader AI safety community sacrifice their principles to appeal to whichever companies are currently in the lead, or something like that.
I think that it’s very plausible that things take into the mid-2030s. I think my median date, my 50% confidence number, is 2038 for transformative AI, which is what I’ve been trying to forecast, which I define as a really big deal. So it’s somewhere towards superintelligence. I’m thinking AI systems that are so capable that if they wanted to, they could take over the world. So it’s the deadline for alignment.
Also that they are moving, say, scientific and technological progress twice as fast as it was prior to that. So if we zoomed out into human history, this would be a time when it’s really happening, as opposed to a time when they’re better than humans at a bunch of things, but there’s only so much compute though, not enough to run a whole lot of copies. I’m thinking of the time when things really are getting going, because I think that the most important role that this plays in people’s thinking is: what’s the deadline for getting impact done by? I think transformative AI is a way to track that.
It could well be, I think, in the middle or late 2030s or beyond. If so, then that’s many presidential administrations away from now, such that it’s very hard to know whether it would be Democrats or Republicans in power. It’s very difficult to know what state America would be in, and whether America is really an ally of Europe and the UK and Australia and other countries, or whether it’s gone in some quite different direction with its threats to invade Greenland and so on recently. That’s quite relevant if one’s thinking about building AI for America or something like that, or questions about should a European alternative be something people are investing in?
If the timelines are two years, then there’s no point trying to do it in Europe, right? But if the timelines are 10 years, it could be the most important thing, or creating a international alternative to a hegemonic national programme.
Rob Wiblin: Ten years is a lot of time for China to change its posture and to potentially catch up. I think even people who are relatively bearish on it.
Toby Ord: Ten years is long enough that the export controls on Chinese chips could well have gone into reverse, where the effective protection that it creates for the native Chinese chip industry may actually have created a big boost to their capacity instead of slowing it down.
It could also be the case that Taiwan has been invaded by that point, and if so, then the West may have lost access to this place that’s systematically producing chips for them rather than for China. It could be a very different situation in a whole lot of different ways. The architectures of the learning algorithms could be very different. It was only, I think, nine years ago that the transformer was developed.
So if we’re projecting another nine years into the future, in 2035, there could be some other major differences where everyone’s using some other weird nonsense word to describe this new architecture that we’d have no idea about. Maybe it has quite different properties to what we’re imagining.
Things could be really different. Another key way is that it’s not clear that there will be companies leading the charge at that point. I think that in general, the longer things go on, the more likely it is that there’ll be national or international projects rather than companies, because people will have realised… if someone just said, “I’m interested in creating the successor to Homo sapiens, that is then going to be more powerful than us in every single way. And not just that, but much more powerful” and so on, we wouldn’t say, “Let’s get a company to do it.” That does seem fundamentally crazy. And that craziness, as we go on, will start to get incorporated into voter preferences and so on, and into national security postures and so forth, where they start to realise we can’t actually let that happen.
Rob Wiblin: Yeah, I guess they’ve been allowed to progress with it because nobody believed that it was actually going to happen.
Toby Ord: Yeah, exactly.
Rob Wiblin: But if they really did think they were about to do it, they would object.
Toby Ord: Exactly. If you have a strategy that’s all about the companies, that strategy could well start to become somewhat moot or misguided as we get into a world where they’re less likely to be the leading actors. There’s a whole lot of things that could end up really different over these longer timeframes.
Another key one is the Overton window. What kinds of policies will be considered reasonable? I think that also connects to the question of how likely is it there could be an international treaty between China and the US? If we’ve seen AI that’s extraordinarily capable, and we’ve seen some of the crazy stuff that’s really done with AI that’s 10 years on from now, it just may not seem at all a stretch that it could take over the world or something else.
So there could be much more appetite for that. If we’ve got double-digit unemployment rates in the West and there’s marches on the streets, general strikes being called for people to stop AI, that completely changes the political landscape — where all of a sudden we’re almost at the situation already to some degree where politicians who just want to wreck AI, where that could actually be a vote-winning possibility, even if it was just totally unconsidered regulation, that all it did was just stick it to these AI people who are immiserating the population.
We could end up in a situation like that where the incentives are totally different. Instead of how do we get as much safety as possible for as little dampening down of the pace of progress, it could be that dampening down the pace of progress is what they’re after. So instead of that being a cost, that’s considered another benefit by the politicians. Things could end up very different, in which case the set of possible interventions could be very different to what they are now.
I think that people by and large are really still thinking on very small margins about how do we tweak the current process by these small amounts. Whereas if it does take a long time, the world just could be so different that the set of available options could be vastly changed.
How should broad timelines change what we do? [02:22:23]
Rob Wiblin: Yeah, people who are already heavily involved in AI, I think they’re very worried about missing the window — doing something that takes three years to pay off and then just everything is obviated by then.
But I guess you’re pointing out that there are some opportunities that might be incredibly impactful that would work if things take eight years or 12 years to pay out, that might be much more useful than the small margins that people are thinking about over the next couple of years. Because if it takes that long, there could be really radical cultural change, opinion change, political change that opens up options that seem completely closed now. It would be, I guess, a bit crazy to just have nobody working on thinking ahead about what those might be or positioning themselves to take advantage of it.
Toby Ord: Exactly. We don’t even quite know what they all are, because there’s so little thought on it. I feel that maybe it would be a good project to even just try to start cataloguing what are the types of things that, if we realised that we had 10 years or 15 years or something, we really would have wished that someone spent the time building these things up.
Rob Wiblin: Yeah, I guess an example is Yoshua Bengio’s project. I guess that feels very sidelined and a bit irrelevant right now. If recursive self-improvement works for Anthropic next year, it probably will feel a little bit irrelevant. But you’re right. Over like four or eight years, if people become gradually more worried, there’s opportunities to get more funding to attract talent, because people think the dominant paradigm is misguided. Yeah, I don’t know how long it would take for that to really pay off.
Toby Ord: I feel that people are maybe having a bit of this issue of anticipated regret. Suppose I decide to write another book. I think a book takes about five years to pay off: from the moment where you have the idea, to you’ve got the publishing deal, to you’ve written the manuscript, to the publishers finally got round to printing it, which is literally about a year — adds up — through to it’s actually had enough impact on the world that it was worth it compared to the short-term things you could be doing, which is like another year or so after the date where it comes out. So I think something like five years.
If I got to the point where I’d written the manuscript, let’s say three and a half years in, and the publishers were getting around to typesetting it and printing it, and then we hit transformative AI, I’d certainly feel like, “Oh my God!”
Rob Wiblin: “I really messed up. I wish I could go back.”
Toby Ord: The anticipated regret would be huge. And that could happen to Yoshua. But if you instead look at the expected value of these things, then I think it actually is like, “We should be taking this.”
Rob Wiblin: A 50% haircut is just actually not that big relative to the differences between projects that already exist.
Toby Ord: Exactly. We need to move out of that, “Is there some chance I feel like a chump?” or something. There’s no textbook on rationality that says that’s how we should make our decisions. But I think that that’s kind of what’s guiding our intuitions a little bit more than that it may turn out that you miss the boat.
But on the other hand, suppose that there’s finally appetite for a treaty to delay the development of superintelligence, but it would require–
Rob Wiblin: Some groundwork to have been laid.
Toby Ord: Yeah, exactly. Either a bunch of diplomatic groundwork or things, but it could also require that there is some kind of alternative approach, some way of getting really high assurance of alignment.
And if Yoshua’s team starts work on that now, such that there’s more chance that they’ve actually got a credible pathway by that point, maybe they find out the initial approach doesn’t work, but they pivot into a second one and that one has legs. If they’ve had time to do that, then maybe they can say there is another option and there’s a reason to have a delay, because if we delay, we’ll still be able to get the fruits of this technology in a safe way. Whereas if they don’t do that work now, then maybe we come to that opportunity and people decide it’s either we get these fruits with a risk or we never get them, and they decide not to delay. There’s a whole lot of things like that.
Rob Wiblin: So the top-upvoted comment on your broad timelines piece, actually on multiple different forums, was this reply from Ryan Greenblatt — former guest of the show, great commentator.
He was like: I agree with basically what you said in this post, except I don’t quite go towards the conclusion.
I agree with many particular points in this post and the apparent thesis, but also think most people should focus on short timelines (contrary to the apparent implication of the post). The reasons why are:
- Short timelines have more leverage. This isn’t just because of more neglectedness now, but also because: (1) it’s easier to target approaches towards shorter timelines where less has changed, (2) short timelines are riskier [and he gives some reasons for that], and (3) it’s easier to operate in “near mode” — [it’s easier to have more concrete thoughts about what to do and what will be useful when targeting short timelines, and it’s maybe also more psychologically healthy to try grappling with the world as it is now]
- I put sufficiently high probability on short timelines: maybe 25% in <2.5 years to full AI R&D automation and 50% in <5.
- I expect work explicitly focused on short timelines (across most areas) to transfer pretty well and generally not cause that much downside in longer timelines.
What do you think?
Toby Ord: My timelines are a bit longer than Ryan’s, but Ryan is characteristically correct in basically everything he says. So I agree with lots of those points. In fact, I’m glad you brought them up because I maybe focused too much on this haircut possibility, that the work you’re doing for longer timelines is obviated by something that happens earlier and makes it moot.
But he’s right that there are additional reasons as well. So that’s not enough in your calculation. There’s also these aspects about additional leverage in the short timelines — due to neglectedness, there’s fewer people working on it; due to some of these other questions about concreteness and so on. For example, if you’re working on AI safety, maybe you’re a bit more likely to go off in some unproductive theoretical direction, or something over longer timelines and things. Yeah, there are a bunch of these additional reasons there.
However, they don’t apply to everyone. Suppose you’re considering a career change, say into government, and what you want to do is work your way up so that you’ll be able to have a senior position on AI in a policymaking capacity. That’s a totally reasonable pathway. And suppose you could get into that position with a high chance or a substantial chance, let’s say 50% chance, you get into that position by 2035, I think it sounds like a very reasonable thing to be doing. And you’ll have a whole lot of short-term goals as to how to work your way up and through that hierarchy and prove that you know things and that you’re a valuable member of the group and so on.
So it doesn’t fall into quite the same things as if you’re a theorist who’s working on AI safety, that maybe like planning against long timelines, you don’t have as much to say. Or if you’re someone like Yoshua, I think he actually has kind of quite a clear, credible idea that he’s going with. If you said, “Drop that and come up with a new plan about what you could do in the short term,” I just think that would be a mistake.
So I think I kind of agree with all of his points, and I think that for certain audiences each one of them are relevant, but they don’t overall say that we should just focus on the shorter timelines.
Rob Wiblin: I guess just different people have very different opportunities available to them. Another one that’s occurred to me recently is, as far as I know, there’s no organisation who takes it as its mission to do research and advocacy and thinking about what are we going to do when there’s no work left for humans to do.
People talk about this in general, but there’s no organisation that has this: figuring out what’s the evidence about when this is going to happen, what should be the government’s reaction? Someone could set that up. It would take many years, I think, to pay off and to build it up into a credible research institute, bring in people who are currently doing that, very scattered. But it would be sad, I think, if no one was willing to do that because they think it will take three years to build up something like a meaningful organisation. Yeah, I guess that’s just one idea of so many.
Toby Ord: Exactly.
Rob Wiblin: Ryan maybe shouldn’t do that because he’s currently doing stuff that’s really useful.
Toby Ord: Exactly. Ryan’s doing the right thing, and is doing very valuable stuff. Even him posting that comment is, I think, valuable. But it just turns out that there are a lot of different pathways that different people are considering. Sometimes they’ve got very strong reason to keep on. Maybe they’ve already spent a year trying to work their way up through this pathway and so on, and that’s given them an unusually high chance of succeeding. They may well be best off continuing to pursue that longer timelines strategy.
Are current models all they’re cracked up to be? [02:31:02]
Rob Wiblin: Let’s talk a bit about how good are the models to actually work with. Are they all that they’re cracked up to be?
On the one hand, I am just super impressed by the models and I use them constantly throughout the workday. I definitely have found, compared to January, I’ve become a little bit less enthusiastic because I think on the occasions where I’ve gone and really scrutinised incredibly closely what the model was doing and what it was saying, it always looks worse.
I guess that probably would be true with humans as well, but I think it’s even more the case with AI is that there’s slippages in reasoning and slippages in evidence, and some degree of confabulation that just makes it really hard to rely on what they’ve said wholesale. It’s a useful input, but not quite as good as what it seemed on the surface when I first started using AIs in this way back in January. Do you have the same impression?
Toby Ord: Yeah, I haven’t run the experiment of really rereading the transcripts. One thing that you probably also notice is you tend not to read every word that the AI produces, and it partly depends on how much… Like when you don’t read some stuff, do you assume it got it right? Do you have a kind of pattern you’re trying to see there of it succeeding? And I think I’m probably more in that direction than the opposite. Although some people see a pattern where they just assume it’s going to fail.
I’ve definitely got a lot of use out of these models as well, including for tasks where I think of them as the following, which is a very kind of academic way of thinking about it, where you’ve just been to a lecture with someone and then you go to the pub afterwards and you’re talking about what you just saw. And maybe the person’s from a different field, and then they kind of like give you some interesting ideas and you ask, “How would that work in engineering?” And they’ll tell you some stuff. Or you just need some advice from someone in a different discipline.
I think that it feels like they’re about as good as asking a researcher at a good university over a beer for a bunch of advice on things — which is to say that when you look back on it, there’s at least one out-and-out error that I can catch, like if I talk to these models for an hour about something.
I was asking this interesting question about how high an orbit around the Earth could you have? Or do you run into issues if you’re trying to orbit somewhere outside the Moon’s orbit, because as the Moon comes around it disrupts you? I was asking about one of these things that was, I think it was five times further out than the Moon. I was asking, does the Moon disrupt this?
And it said, from this distance, the Moon has more gravitational pull than the Earth. And I was like, that’s clearly not true because the Moon’s like a sixth of the diameter of the Earth, and you’re five times further out than the Moon’s orbit and the Earth’s just so much bigger. And I was like, “Are you sure that’s right?” And it was like, “Oh, no, actually it’s a thirtieth as large. I overstated that.” It’s like, you didn’t just overstate it — you’re out-and-out wrong. So it’s funny, I was talking to a colleague about this and saying I do find that, if I talk to it about something, there’ll be something like that where even a nonexpert in that area is like, “Hang on, what? That doesn’t sound right at all.”
Then you check and it’s like, oh, no, it wasn’t right. And his response was, “There’s more times than that, you just can’t see them, because you’re asking about an area that’s outside your discipline.” There’s some smallish number of things that even someone from outside of the discipline can catch. But there are more mistakes than that, which is a bit alarming.
Rob Wiblin: Especially given how much we’re learning to rely on them. Just from a prudential point of view.
Toby Ord: Yeah, so maybe it’s not quite as good as the colleague who is a bit glib at some point, and they eventually say, “I kind of assumed that you wanted me to assume such and such. I didn’t realise that you didn’t. So I guess my statement wasn’t really right,” because even a human colleague will say things that are wrong. Partly because it’s just an out-and-out blunder, and partly because they had a slightly different model of what question you were asking and they answered the wrong question or something like that. But yeah, I think this made me realise that it could be actually a little bit worse than that still.
Rob Wiblin: Yeah. I’ve also started to think of them as having the goal of having a glossy report that you will like, that will seem good. I think they reason about this in their chain of thought, they think about you when they think about what will appeal to you. I think we’ve managed to tamp down on the outright sycophancy where they’re just absolutely blatantly flattering you. But I think it’s a bit like a student that’s trying to make an essay that sounds good to the teacher who’s just scanning over it. That’s another way in which they’re a bit weaker than you might think, if you just started playing with them briefly.
I think probably Fable is less bad in this regard. Not on the glossy thing, but I think probably it actually is just a significant step up. But I guess I am always doing this adjustment now. Well, I don’t know how far to adjust. It’s such a difficult thing, that they’re somewhat less impressive than they seem, but is that going to be a persistent thing, or is this just like a minor downgrade? Or is it like a more fundamental problem that they don’t have a grasp of reality? They’re just kind of speaking.
Toby Ord: Yeah, maybe all of these things. Sometimes we still see examples of them really making serious mistakes. I saw an example recently which is like, “If an English person says to an American, ‘What would happen if you lift your coffee cup?’ Then what would happen?” And it was like, “Oh, the classic confusion because the American will think it’s about an elevator.” It’s like, what? No, they won’t.
Occasionally there’s just stuff like that, but that’s generally restricted to the really small models, like the one that happens if you type something to the search box in Google or something. So we see less of those word-salady kind, like just what the hell happened there? But there could be still a bit of that.
But I like your example about how they’ve kind of got a model of you and they’ve been trained to say things that you will find convincing, which is not nice. That’s how a lot of people write. You get taught that in school how to write a persuasive essay where effectively you’ve got a model of the subject and you’re trying to convince them of something. I don’t want that.
Rob Wiblin: It’s a bit adversarial.
Toby Ord: Yeah, it is adversarial. I don’t want my AI to be tempted to convince me of stuff. I would like it to be laying out the facts in a way, frankly, that points to the greatest weaknesses and the greatest strengths of the argument that it’s just laid out and so on. And I don’t think it’s good at all that it’s attempting to sell me on something. I think that is an area where we’re going to overestimate it.
Another one is that — when it comes to tasks without verifiable rewards — and we try to think, how do they train these tasks? One of the ways they hope to get better at that is by transfer from tasks with verifiable rewards. There has been a little bit of that, where it learns some reasoning technique like saying, “Hang on, wait, let me check that.” Then it learns that from the verifiable area, like mathematics, and then it starts using those kinds of tricks when it’s reasoning about the nonverifiable ones. That’s good. I’m not sure there’s that much more transfer going on there.
But there’s also attempts to train them with these LLM-as-judge approaches, where there’s no fully verifiable way to sort out how good its answer was, but you could have another language model assess the answer and then give a rating and so on. And if the language model starts off better at assessing the quality of things than it is creating things, it can kind of piggyback on its discriminative capabilities and turn those into generative capabilities.
So you might be quite good at working out whether a novel was written well or not compared to how good you’d be at writing a novel. If so, you could use one of these techniques to become better at writing a novel, where you write a whole lot of novels and then you kind of assess them.
Rob Wiblin: “I should do it more like this.”
Toby Ord: Then you kind of slowly learn and adjust. That’s a clever technique and it’s produced some value in these nonverifiable areas. But if you think about what goes on there, what does it train for? It trains for things it can detect, and in particular, if there’s an out-and-out mistake that it can put its finger on, it teaches it not to do that. I’ve certainly noticed that the outputs of language models on open-ended questions and things are getting harder to point to the fact it’s made an out-and-out mistake. But are they actually getting more accurate, or are they getting better at going into some unfalsifiable direction where it’s harder to pinpoint that it’s made a mistake? It’s not clear.
If you look at the theory of what you’d expect, yeah, I guess I would expect them to be getting a little bit obfuscated and so on.
Rob Wiblin: Basically making it hard to assess the reasoning very clearly. Making it harder to spot your own mistakes, I suppose. Or vague enough that errors are not apparent. I think, yeah, I guess you’re right. Theory would predict that they are becoming like that. And I perceive them as becoming like that.
Toby Ord: Yeah, I think so. And I want to go back right at the start of this. You said they’re impressive, I just want to stress that. But they are genuinely impressive. And occasionally they do things, clearly achieve things, which are kind of impressive and one shouldn’t lose track of that. But by the same token, when you do reexamine it, yeah, they’re maybe not as good as you thought.
Rob Wiblin: How much of a discount is this on prospects of getting recursive self-improvement or AGI? It’s quite hard to assess, because if it is just a thing of like, well they’re 20% worse than you think, then we’ll just wait three months and then they’ll be as good. You know what I mean? Because they’re improving so much, this only creates like a small delay basically.
If it’s something more fundamental, that they’re heading in the wrong direction, they’re trending in the wrong direction, and this thing isn’t getting fixed because we don’t have ways of getting them to be more precise in this respect—
Toby Ord: Yeah, it could be they’re trending in the wrong direction. It could be that they’re actually plateauing at some kind of intermediate quality that they can’t get beyond. That would also be a problem.
I think that really a lot of the recursive self-improvement question comes down to this question of: is AI fundamentally hill climbing in a way that you can do with relatively small amounts of insight that doesn’t involve anything like… if we look at the history of AI, there have been a bunch of really big and original ideas that I would call creative.
The idea of connectionism, where we should be, instead of just thinking about logic and the structure of first-order logic and reasoning, we should be thinking about a whole lot of little dots with little lines connecting to them, with concepts connected to other concepts and so on — which led to the perceptron, which led to the neural network. That idea that we should throw away all of these logical formulas and instead think about these dots with arrows connecting them and so on, that was a pretty big and creative idea, obviously inspired by the brain.
Another one is the idea that you can think of intelligence as compression, that fundamentally to really understand something means that you’d be good at expressing it in as small a way as possible. It’s a very deep idea and it’s really not obvious.
It’s not clear that the current AI systems could come up with stuff like that. So are we in a world where you need another breakthrough at that kind of level, that kind of creativity, with the kind of thing that comes around like once a decade or something for the whole research community of AI scientists? If so, there could be a bottleneck. And our perceiving of looking at these transcripts and noticing that they’re slightly tricking us and things would be pretty bad news for them.
But if it does turn out that we’ve got all of those that we needed, all we need to do now is scale and maybe there’ll be a more efficient way to get there with one of those things, but with enough scale we can just get there anyway. If we’re in that world, then that’s the kind of world where the RSI may well be quite likely.
Coordinating careers for different timelines [02:43:35]
Rob Wiblin: To close out, coming back to the broad-timelines philosophy, do you think people in the audience — if they want to be working on making AI go better — to what extent does it imply that each person should be thinking, “I want to do something that is useful across many different possible timelines, something that’s useful if it comes in 2030, 2035, 2040?”
Or does it more imply that we should have a portfolio across lots of different people, and people should spread out working on different projects that pay out optimally at different points in time? Do you have an idea?
Toby Ord: I think it’s quite a lot more like the second there. So as opposed to that everyone should just be imagining they’re the only actor and they have to address all of these different possibilities — maybe a government should be thinking like that, they need to have policies that would kind of address these different possibilities. But for an individual, I don’t think that’s right.
I think that the main thing is that they shouldn’t be acting as if there’s a particular certain timeline that they have to get all their work done by. Instead, there should be a bit of a discount, like a lowering of the value of things that would pay off over longer timelines compared to what you’d otherwise think, and a bit of a hedging towards doing things that hedge against these earlier possibilities.
But ultimately you really do see it at this portfolio level that what we want is that the people who are really good at running in the marathon, the people who have these really great opportunities for building something much larger than themselves, where they think that they could have 100 times the impact if they were to take eight years establishing this thing, we want them to be doing that. We don’t want people who are focused on short things to switch into that. And then we want the people who are really well situated to do work now to keep doing it.
Then ideally some people would take a step back from that and see: are we slightly overindexed on one of these or the other, where maybe we should direct some people who have equally good options as to where they go?
Rob Wiblin: My guest today has been Toby Ord. Thanks so much for coming back on The 80,000 Hours Podcast, Toby.
Toby Ord: It was wonderful. I’m sure I’ll be here again.