#250 – Toby Ord on where AGI timelines go wrong

Both Silicon Valley and the public can’t get enough of ‘AGI timelines.’ But Toby Ord, senior researcher at Oxford’s AI Governance Initiative and author of The Precipice, believes we consistently make big mistakes when thinking about them. He lays out the 14 ways he most often sees people go wrong:

  1. Assuming AI research is just hill-climbing
  2. Imagining AI research is just programming
  3. Forecasting “could” instead of “will”
  4. Believing the current benchmark is the last one
  5. Extrapolating trends with no clear finish line
  6. Assuming inputs keep scaling at the same rate
  7. Conflating intelligence with capability
  8. Consuming point estimates and discarding the error bars
  9. Dismissing dissenting experts
  10. Forecasting very different things while using the same words
  11. Assuming capabilities arrive together
  12. Treating “we don’t know” as permission to carry on as usual
  13. Choosing a plan that minimises regret rather than maximises impact
  14. Trusting surface model impressiveness

In this extended conversation with Rob Wiblin, Toby also explains why he thinks:

  • AI self-improvement is uniquely dangerous in four ways, but also might not even work
  • A ban on superintelligence is possible
  • A US-China treaty on superintelligence is also possible
  • The case for ‘broad timelines’
  • Transformative AI is likely a decade away
  • We should just ban unmonitorable chain-of-thought today.

This episode was recorded on July 2, 2026.

Want to get up to speed on AI?

We’ve got a crash course of 10 of our podcast episodes designed to help you get to grips with transformative AI — particularly if you’re new to the topic — and what you can do to help shape its trajectory.

Our production team includes:

  • Video editors: Josh Alward, Dominic Armstrong, Ollie Bignell, Andrés Escobar, Jasper Luithlen, Milo McGuire, Luke Monsour, and Simon Monsour
  • Producers: Elizabeth Cox and Nick Stockton
  • Coordination and support: Katy Moore and Lou Moran
  • Camera operator: Jeremy Chevillotte
  • Music: CORBIT

The episode in a nutshell

Toby Ord — senior researcher at Oxford University’s AI Governance Initiative and author of The Precipice: Existential Risk and the Future of Humanity — thinks recursive self-improvement (RSI) probably won’t ignite an intelligence explosion, and artificial general intelligence (AGI) could easily be more than a decade away.

But “probably” is the key word: he still finds an explosion credible, and even without one, the possible speedups from RSI make it one of the biggest issues in AI.

He argues this is reason enough to prepare now — with a moratorium on superintelligence, a verifiable US–China deal, moral norms against the most reckless research, and a portfolio of work robust to AGI arriving both next year or after 2040.

RSI (probably) won’t cause an intelligence explosion

Toby distinguishes today’s version of RSI from the picture developed by Eliezer Yudkowsky:

  • In the older picture, a single program repeatedly rewrites its own source code, rocketing to superintelligence before humans can intervene.
  • Today’s version is a gradual feedback loop: thousands of researchers and AI systems working together, with AI handling coding and simple experiments first, then potentially more of the research process over time.
  • The difference matters because coding may not be the bottleneck. AI research also means deciding which experiments to run, spotting the important conceptual problems, and inventing new architectures — strategic abilities current systems haven’t demonstrated.
  • Even if AI could replace the full range of human research skills, progress might simply flatten off against rising complexity or fundamental limits, rather than bending upwards into an explosion.
    This gradual version looks somewhat safer: humans can stay involved, give feedback when things go wrong, and monitor systems with interpretability tools — rather than having to specify human values perfectly before pressing return on an irreversible process.

RSI still creates at least four distinct kinds of danger

Even if RSI only produces a temporary acceleration, Toby sees four ways it could make things substantially riskier:

  • Capabilities could outrun safety:
    • AI may accelerate programming and capabilities research more than it accelerates alignment or governance.
    • At any given capability level, we’d then arrive less prepared and with weaker control mechanisms than we otherwise would have.
  • Misalignment could propagate through generations:
    • If GPT-8 was scheming while designing GPT-9, it could preserve its objectives in its successor.
    • Therefore, a single misaligned model could compromise an entire lineage of increasingly powerful systems.
  • Society could lose its warning shots:
    • With cars, guns, and other technologies, people encountered weaker versions first and learned from manageable harms.
    • A sudden jump equivalent to five years of progress in one year could skip the intermediate failures that would otherwise provoke regulation or safety improvements.
  • The first mover could gain an overwhelming lead:
    • When progress is steady, a lab that starts one year behind stays roughly one year behind.
    • When progress bends sharply upwards, the first lab to trigger RSI could become many effective years ahead — increasing racing pressure, concentrating power, and making the outcome more winner-takes-all.

US–China agreement is more plausible than people assume

Toby believes the US and China could eventually agree not to build superintelligence (though an immediate deal governing RSI may be harder).

His central argument: today’s scepticism assumes the dangers will stay as abstract and controversial. But if progress is gradual, increasingly capable systems will make the threat visible:

  • People currently have to extrapolate to see why advanced AI could be dangerous. Closer to the threshold, they’ll only need to look at what existing systems can already do.
  • Once leaders believe their own families, countries, and cultures are genuinely threatened, ordinary self-preservation may do more work than sophisticated arguments about existential risk.

He also thinks the proposed alternative — racing to develop superintelligence first and using it to dominate or disarm the other side — isn’t strategically viable:

  • If the US were approaching a technology that could neutralise China’s nuclear deterrent or force regime change, China would have a strong incentive for a preemptive strike — similar to a missile-defence system that renders an opponent’s nuclear arsenal useless.
  • Strategies based on “getting there first” often fail to consider the adversary’s best response if this appears to be happening.

The key obstacle to an agreement is verification, but Toby thinks people dismiss it too quickly. Options could include location-aware chips, physical inspections, or matched reductions in compute.

A moratorium would be better than a pause

Toby prefers a moratorium to a pause — a pause implies development automatically resumes, whereas a moratorium means society has to decide there’s a compelling case for reversing the ban.
He doesn’t yet know exactly where a moratorium should draw the line. Options include a capability ceiling, a speed limit on annual capability gains, or compute and token limits on using AI for AI R&D.

And humanity has established neither the safety nor the legitimacy needed to create superintelligence. Toby believes:

  • Ordinary evidentiary standards should apply when imposing enormous risks on strangers.
  • The decision shouldn’t be made by Silicon Valley — or even the West — before most of the world has seriously considered it. If less than half of humanity thinks superintelligence is desirable, building it would be extraordinarily hard to justify.
  • Even a 1% chance of extinction is roughly 80 million expected deaths — larger than World War II, before counting future generations — yet the effort devoted to reducing the risk is nowhere near commensurate.

A moratorium could backfire — provoking backlash, letting compute accumulate, or leaving many actors crowded just below the threshold, ready to race when it lifts.

But the opposite is also possible: a long-running moratorium could become normalised, discourage speculative investment, and grow easier to maintain over time. And if compute accumulation is the worry, restrictions could extend upstream to the handful of companies producing crucial chipmaking equipment.

Moral norm changes shouldn’t be underestimated

Toby thinks AI discourse underestimates moral change. Analysts assume people and companies will inevitably follow financial and strategic incentives, but societies have developed strong norms against things they’re perfectly able to do:

  • Human cloning and eugenic enhancement aren’t avoided because every laboratory is monitored, but because most researchers don’t want to do them.
  • Women’s suffrage, environmentalism, and other moral changes can’t be explained well by modelling everyone as narrowly self-interested.
    If researchers come to regard a line of AI research as profoundly wrong — and its creators as potential villains of history — persuading people not to want a technology may be more tractable than constructing incentives strong enough to stop those determined to build it.

Unmonitorable chain of thought is a great example. Toby thinks research that makes AI reasoning illegible while delivering a meaningful capability advantage could be among the most harmful ever conducted.

Legible chain of thought is an extraordinary, unexpected safety advantage:

  • Earlier thinkers would find it miraculous that AIs reason in human language, letting researchers essentially read a model’s thoughts. Throwing that away for a modest efficiency gain would be a terrible trade.
  • Labs already appear reluctant to train away legibility, so the beginnings of a norm may already exist — this could be strengthened by cross-company agreements, research standards, and whistleblower protections. The longer the norm survives, the harder it is to violate.

Toby’s timelines: median 2038, but “we don’t know”

Toby offers several reasons AGI could remain more than a decade away:

  • Benchmark progress does not reveal how many benchmarks remain:
    • AI repeatedly goes from weak performance to saturating a benchmark.
    • But there may be one more layer of capability to master, or ten.
    • Even if METR’s task-length trend extrapolates perfectly, nobody knows what task horizon corresponds to AGI.
  • Scaling may slow down:
    • Early compute growth partly came from companies spending more money and redirecting existing chip capacity towards AI.
    • Those sources of growth cannot continue indefinitely: manufacturing capacity takes time to build, and the rate of successive 10-fold increases may fall.
    • Human-generated training data is limited; one new year of books and scientific papers contributes little relative to the accumulated corpus.
  • Recent leaps are concentrated in tasks with verifiable rewards:
    • Reasoning models transformed maths and coding because training could loop — attempting the problem repeatedly, checking the answer, and reinforcing what worked.
    • Almost everything else offers no such check: an AI learning to run a cafe profitably would have to run 100,000 unprofitable ones first.
    • For these non-verifiable skills, there’s little theory about how systems reach superhuman level — which explains why capabilities are so jagged.
  • Further conceptual breakthroughs may be necessary:
    • If scaling hasn’t delivered extremely powerful AI by the early or mid-2030s, it may indicate that AGI can’t be reached through algorithmic improvements alone.
      Toby advocates broad timelines because expert disagreement and forecasters’ probability distributions are extremely wide. His own 80% confidence interval is roughly three years to beyond 100 years.

His median for transformative AI — capable of taking over the world, at least doubling scientific progress, representing the deadline for alignment — is 2038. But he can’t rule it out within nine months; he just thinks it’s less likely.

Toby stresses broad uncertainty is not permission to choose whichever timeline is most convenient. Governments should hedge against early AI while also building institutions, movements, treaties, and technical approaches that may require many years.

Superintelligence might not be job-ready

Even if RSI created an extraordinarily intelligent system inside a data centre, Toby argues it might lack the capabilities to transform the economy immediately. He compares it to an exceptionally bright university graduate:

  • It might have greater raw intelligence and potential than any human — but it would still be applying for entry-level roles, lacking practical experience, tacit knowledge, relationships, and trust.
  • Running Apple or being a law firm partner requires understanding a culture and having lived through unusual situations, mistakes, conflicts, and business cycles — knowledge that may require actual calendar time, particularly for once-a-decade events.

The lag between superintelligence and senior-level competence could be anywhere from six months to 20 years, depending on the job. Pooled experience across many copies, better sample efficiency, simulations, or new workplace data could shorten it — but no number of copies can experience events that haven’t happened yet.

AI assistance may also produce less human learning than the productivity figures suggest: a researcher who delegates a four-hour task misses the understanding they’d have gained doing it — and nothing speeds up the offline processing during sleep and breaks that lies behind many scientific breakthroughs.

AI disproved a famous conjecture — but maths is more than proofs

Toby considers the AI disproof of Erdős’s unit-distance conjecture genuinely impressive. It connected two areas of mathematics not previously known to be related, giving mathematicians new ideas to import. But mathematicians’ fears of obsolescence focus too narrowly on proving formal statements:

  • There are infinite mathematical claims, so the first challenge is choosing which questions are interesting — and higher still is inventing new concepts and branches, like Claude Shannon did with information theory. Current AIs haven’t shown that creativity.
  • Calculators automated mathematicians’ numerical work and symbolic software automated algebra — without ending mathematics. So automated theorem proving could instead push mathematicians towards choosing questions and creating theories.
  • These abilities might also be automated at some point — but history repeatedly reveals another capability layer after the supposed “final rung” has fallen.

Today’s models are helpful, but wrong more than you think

Toby finds frontier models useful — like discussing an unfamiliar topic with a good university researcher over a beer — but expects at least one blatant error per hour of conversation, plus more mistakes that a nonspecialist can’t detect.

Part of the problem is what they’re optimised for. They have a model of the user, and produce answers the user will find convincing — the equivalent of a persuasive essay, which Toby considers adversarial.

He’d prefer an AI that lays out the facts and frankly points to the greatest weaknesses of its own argument, rather than one trying to sell him something. Models are now better at avoiding detectable mistakes, so answers may be getting vaguer and harder to falsify rather than more accurate.

This matters for RSI: if models are less capable than they appear, how much does that discount the chances of AI driving its own improvement? For Toby, it largely comes down to whether progress mostly requires hill climbing — which current models may manage — or rare creative ideas like connectionism or intelligence-as-compression, which they haven’t yet shown they can produce.

Broad timelines call for a balanced portfolio of approaches

Toby agrees short-timelines work has special advantages — more concrete, more neglected, and targeted at the period when society is least prepared — and thinks governments still place too little weight on transformative AI arriving soon.

But he rejects blanket advice to work only on projects paying off within a year or two:

  • A five-year project might lose 20% or even 50% of its expected value because AI could arrive first — easily outweighed if the project would otherwise have 10 or 100 times the impact.
  • Organisations, intellectual fields, movements, and senior government careers take years to build. Worthwhile projects include international treaties, alternatives to current alignment methods, or institutions focused on a world with little work left for humans.
  • The longer things take, the likelier national or international projects are to displace companies at the frontier, and mass unemployment could make slowing AI a vote-winner — this could open options that look closed today.
  • People anticipate the regret of years spent on a project suddenly rendered obsolete — but feeling like a “chump” isn’t the same as having made the wrong decision in expectation.

His preferred solution is specialisation across a portfolio:

  • Governments and the safety community should ensure somebody is addressing each plausible timeframe, but individuals don’t each need to hedge every timeline.
  • People placed to achieve something now should keep at it, while those with rare opportunities to build institutions or rise to influence over a decade shouldn’t abandon them because AI might arrive first.

The mistake is not choosing a particular horizon — it is acting as though it were certain, and letting everyone else make the same bet.

Highlights

Nuclear disarmament shows mutual compute-reduction verification is possible

Toby Ord: I think there’s a bunch of ways to do [verification for a US-China agreement], and it depends on exactly what the technology is at the time and so on.

People have often been thinking about some very technologically advanced verification mechanisms, such as having the next version of the NVIDIA chips be able to know exactly where they are in the world with GPS. And if someone blocks the GPS, then maybe the chip shuts down. It will only start working again if it can share where it is, and so on. I mean, that would be great if we get those things. Then, for example, there would be credible ways to prove that your chips aren’t actually working on advancing AI further.

But even I think people forget that the universe of possible ways of doing verification is very large. And I think the Cold War is really illustrative. For example, one of the key things was that the US and the USSR wanted to reduce the number of nuclear-capable bombers that could threaten the other person’s home soil.

Eventually some bright person, thinking outside the box, worked out that these bombers were on the runway, and they could saw them in half with a huge circular saw. And then, with tractors, pull the halves apart from each other so that spy planes from the opposing side could witness this and credibly see the destructive capacity be destroyed. As far as I understand, those planes are still on the runways, in halves, because someone was like, “Hang on, this is a way you could do it.”

That’s the kind of thing you come up with after you’ve been thinking about this for quite a long time. Whereas, at the moment, there’s a lot of people who’ve thought about it for less than one day, like less than eight hours full-time thinking whether verification is possible, who say, “It’s probably impossible, and so I guess we will all die.” I feel that is a bit of a ridiculous approach.

It’s reasonable to think it might be impossible to do the verification. Maybe we’ll think about it for years and we won’t find any ways. But I think that there are a lot of credible ways, and that’s an example that you couldn’t have initially done. It required the Overton window to move a bit before that one was possible — for the US to say, this will literally involve destroying capacity we’ve already built — and that would have been annoying at first, but once they realised that it would also destroy the same amount of capacity on the other side, you start to think even though we’ve already built it, it’s probably worth doing.

I think an example that would be very analogous to that is: at the moment, there’s a lot of chips that aren’t being tracked, and where we don’t know exactly how many have been smuggled into China and things like that, but you could have, say, like-for-like destruction of that capacity.

Moral norm changes are an underrated intervention

Toby Ord: If you can make it so they don’t want to do [unmonitorable chain of thought reasoning], if you can win the moral argument — that it’s bad, you’re a bad person if you do this behaviour; possibly one of the worst people who’s ever lived actually, if you create this technology — that changes things.

If you ask why isn’t anyone cloning any humans, or why isn’t there a whole lot of genetic engineering going on creating kind of genetically superior humans? If you went back and looked at what everyone was saying from 1880–1920 or something, every intellectual group was talking about eugenics and it was just generally thought that, once the capabilities were there to understand how to actually change the makeup of people in future generations, that everyone would be doing it.

And I think they would be pretty bewildered to find out that we actually gained the abilities to do that and then just no one’s doing it. And they’d say, “Is that because you’ve got all these verification techniques and you’re in the labs in all these different countries just making sure that no one’s doing it?” And it’s like, “No, we did this thing–”

Rob Wiblin: Nobody wants to.

Toby Ord: We stopped wanting to do it. It’s like, well, how did you do that? And the answer is moral change and norms change.

People really undervalue this. I think that this is a different approach to trying to use narrow incentives, to try to resist the massive corporate incentives to keep building more and more powerful systems. That’s a really hard battle to win. But actually changing the norm so they no longer want to do something can definitely work.

Rob Wiblin: Yeah, you’re so idealistic, Toby. I feel like this kind of idealism is rare in AI discourse at the moment. Or the idea that well, we could just, for moral reasons, not do stuff. I don’t know. But there’s this analytical frame that you get into where it’s like, “Well everyone is going to just follow their incentives to do that.”

Toby Ord: Yeah. I think it’s rare in the world in general to think this way, but I think it’s powerful. And I feel like if you were an economist and you’re trying to understand, say, women’s suffrage in England back when those debates were happening, you would have thought there’s this currently kind of privileged class of people — men, or if you go back further, landowning men — and why would they share their franchise with other people? Currently they’re holding almost all the power. Why would they ever do this?

And I think that something along the lines of “it’s the right thing to do” is actually a key aspect of what made those things happen…

Rob Wiblin: I feel that someone who is just predicting everyone’s behaviours as like narrowly self-interested, kind of psychopathic people in these economic models, I think they would have made bad predictions.

What they do is then eventually, maybe in their utility function, one of the things they care about is like their wife, or one of the things they care about is justice, or maybe justice gives them a warm glow which is one of the forms of happiness. They can have a Mars bar or they could have some justice. They have these ways of trying to incorporate it in, in order to make sense of something after the fact, but they rarely think of that in prospect when they’re trying to predict whether something will happen.

Toby Ord: But it turns out that, “Why didn’t I do some thing? Because I thought it was deeply wrong” is often the answer. Also that we can change what we think is wrong and we can learn that something’s wrong.

Toby's broad timelines and their implications for impact

Toby Ord: I’m also very worried about short timelines. Thinking that it may well be long timelines doesn’t make it all that much less worrying that it also may be short timelines. I think that you’re right that there’s a bunch of ways that it could come soon, especially if recursive self-improvement happens and if it’s at the easy end of the spectrum.

We don’t know whether the field of AI, in order to get to these really advanced levels, if it’s mainly just hill climbing and making small tweaks to the current basic structure and then just seeing what makes it go better and just following that gradient. Could be, and if so, then there could be really explosive growth in AI capabilities. We can’t rule out, say, Daniel Kokotajlo’s nine months. I don’t think I can rule out that it could happen in nine months. In fact, I can’t really rule out that it’s already happened behind closed doors and I haven’t heard about it yet, but I don’t think that’s likely.

What I try to say to myself with some of these things is it could well be that we have, say, transformative AI before the end of the current presidential administration in America, but we probably won’t. And so it’s useful. It’s important to know that it could happen. It could be the current political arrangements are the arrangements under which it occurs. But I think that it’s more likely that they’re not, in which case things could be really quite different.

So that’s an example of how you can learn from this perspective that you need to be able to hedge against these early possibilities that happen when we’re least prepared, but not overcommit on them or something.

I worry if companies sacrifice their principles in order to appeal to whoever’s currently in charge, or if people in the broader AI safety community sacrifice their principles to appeal to whichever companies are currently in the lead, or something like that.

I think that it’s very plausible that things take into the mid-2030s. I think my median date, my 50% confidence number, is 2038 for transformative AI, which is what I’ve been trying to forecast, which I define as a really big deal. So it’s somewhere towards superintelligence. I’m thinking AI systems that are so capable that if they wanted to, they could take over the world. So it’s the deadline for alignment.

Also that they are moving, say, scientific and technological progress twice as fast as it was prior to that. So if we zoomed out into human history, this would be a time when it’s really happening, as opposed to a time when they’re better than humans at a bunch of things, but there’s only so much compute though, not enough to run a whole lot of copies. I’m thinking of the time when things really are getting going, because I think that the most important role that this plays in people’s thinking is: what’s the deadline for getting impact done by? I think transformative AI is a way to track that.

It could well be, I think, in the middle or late 2030s or beyond. If so, then that’s many presidential administrations away from now, such that it’s very hard to know whether it would be Democrats or Republicans in power. It’s very difficult to know what state America would be in, and whether America is really an ally of Europe and the UK and Australia and other countries, or whether it’s gone in some quite different direction with its threats to invade Greenland and so on recently. That’s quite relevant if one’s thinking about building AI for America or something like that, or questions about should a European alternative be something people are investing in?

If the timelines are two years, then there’s no point trying to do it in Europe, right? But if the timelines are 10 years, it could be the most important thing, or creating a kind of international alternative to a hegemonic national programme.

Superintelligence might start out as a brilliant grad student

Toby Ord: Maybe you could reach a kind of superintelligent system in a data centre, where it’s training and learning a bunch of things. It gets really skilled at things like mathematics, and maybe by self-improving it gets really good at a whole lot of other things.

Suppose you did. Suppose there was no obstacle to how smart it was. And so the thing that came out of this data centre at the end of this, you could think of it like an amazingly bright student who’s just finished their undergrad degree; they’re going to go out into the workforce, and the world is their oyster. They could go into politics, they could go into business, into tech, into finance, into journalism. And let’s suppose they’re so capable, they’re the most capable ‘person’ applying for the job in any of these areas, but they would be applying for an entry-level job. So the most capable person with the most potential for journalism who hasn’t yet—

Rob Wiblin: Ever written anything.

Toby Ord: Exactly, exactly. They haven’t yet reported on a hot-button issue and then got a whole lot of flak or whatever and had to work out how to navigate it, and they haven’t worked out how to trade off the risk to the reputation of the newspaper they’re writing for vs the integrity to the facts and so on.

I think what I would say is, even if you could train something with recursive self-improvement to become extraordinarily intelligent, there may just be a lot of skills and capabilities that it won’t yet possess. And you wouldn’t fault its intelligence for that. You wouldn’t say it’s less intelligent because it doesn’t possess these skills. Suppose, for example, Tim Cook is stepping down from the CEO of Apple and they’ve already got someone lined up, but suppose that they didn’t and they thought this AI is superintelligent, we could have it be the next CEO of Apple. Well, I’m not sure about that. In the same way as if there’s a really bright student who’s just finished college.

Rob Wiblin: The most fast-learning 21-year-old, I guess.

Toby Ord: Yeah. You’d still say no, actually there’s a bunch of experience about the culture of Apple, for example, that you would need to know in order to do this job.

And also a lot of relationships and trust that some people have built up. It helps you realise that it’s connected to the diffusion question of how much of a lag is there from having these amazing AI capabilities demonstrated to them, to actually being out everywhere in the world? Where it could be that to be a great CEO of a Fortune 500 company, that it does actually take a decade or more of just calendar time in terms of approaching new opportunities, seeing the business cycle happen, seeing people who are betting the wrong way get wiped out by the market, and it’s not just cases where you can learn from what’s happened in the past. There’ll be a whole lot of new conditions that have never existed before, including with AI totally changing everything.

So it made me realise that, instead of saying you can’t become superintelligent in a data centre, I’d say even if you become superintelligent in a data centre, that doesn’t necessarily make you super capable and able to do all of these jobs. What it would probably leave you with is being a great entry-level employee in anything, and advancing along their career trajectory faster than a normal human would. So you’re better than a relatively fresh-faced 21-year-old of any stripe or something, but not that you’re better than people with 30 years of hard-won experience or context on those jobs.

That could mean that that doesn’t say much about how transformed will the world be in 50 years’ time, but it does say that actually maybe if you had this amazingly intelligent AI in a particular year, maybe a couple of years later, the world isn’t that transformed because it can’t be doing most of the jobs yet.

So that did make me think about an additional delay to add into this kind of calculation of: at which point do you get recursive self-improvement, if it’s possible? Then how long does it take to reach a highly intelligent system? OK, but then there’s another delay between how long does it take to have the highly intelligent system, let’s say, and it getting into a whole lot of companies.

Rob Wiblin: And having lots of concrete skills.

Toby Ord: Then once it’s in those companies, how long does it take from entering that company to understanding enough about these types of roles in order to be able to really deliver at a senior level of performance? It may add, I don’t know, somewhere between six months and 20 years, depending on the particular job of those areas.

Automated proofs could enrich mathematicians' jobs, not destroy them

Toby Ord: There’s infinitely many true mathematical statements. So you have to kind of greatly prune this space of statements to the interesting statements.

One way to rephrase that is like, what questions should we be asking? What are the interesting mathematical questions?

And it’s not clear that [AI] can do that, that it can work out which questions to ask. At that point you’d still need humans to say, “This one is one of the ones I want to actually have the system spend its time on, not this exponentially growing thicket of uninteresting questions.”

So that’s the first step. Then beyond that there’s even richer things, there’s this question of: the AI systems can take some formal statement in some theory of mathematics — in this case discrete geometry — and then try to prove it, but they can’t create new theories of mathematics.

So if you go back in time, say before Claude Shannon invented information theory: now, today, we could ask questions about the information that can be carried over a noisy channel and optimal coding and stuff like that. But we didn’t even know how to ask those questions back then. We didn’t know what kind of formal statements to have that would correspond to these kinds of informal and inchoate ideas that we hadn’t fully pinned down.

So when mathematicians invent areas like that, there’s a kind of extreme creativity. If you go back to the 19th century, there’d been thousands of years of geometry, and then in the 19th century mathematicians worked out that you could ask questions about things beyond three dimensions, like a four-dimensional cube: how many corners would a four-dimensional cube have? But they didn’t ask these questions prior to then. They also worked out, in the 19th century, you could have curved spaces, and they didn’t realise that we actually live in one, but they just thought it was an interesting mathematical question. They also kind of worked out about fractional dimensions, like things between two and three dimensions.

And there was this burst of creativity that you could ask all of these types of things. Or if you think about the origin of calculus, that we could ask about not just how high up is some curve, but we could ask about the slope of the curve and how that changes over time and so on. And that once they came up with these theories, all of a sudden there was this whole infinite range of new interesting questions you could ask. And the AI systems haven’t shown that they can do that at all.

So I think it’s instructive to see that, even if mathematicians spend a lot of time proving things, there’s these other layers that haven’t really begun to be automated.

Rob Wiblin: So they can’t do it now. But I would expect that stuff is coming soon. Do you agree?

Toby Ord: It might be, yeah. I’m not claiming that it won’t be able to do it, just that there’s this kind of thing that we think we can see this trend. We’re laser focused on this issue about proving things or something. Then we think that a mathematician, that’s what they do, and that it’s going to automate that away. But it’s easy to lose track of the fact that there are these higher-level questions, which I’ve always thought are the more important questions in mathematics. I’m less impressed if someone proves a difficult theorem than if they invent a new branch of mathematics where entirely new questions come into view and new concepts. That’s always been what I thought was the more impressive thing.

If we look at the history of automation of mathematics, for most of the time — until the 20th century actually — mathematicians spent something like half their time doing numerical calculations. Then the calculator automated all of that away, so it automated half of a mathematician’s job. And we don’t think that was a great shame or something. We also didn’t think we were on the verge of a singularity or something when we did that.

And then from about 1980 to now, symbolic manipulation, like solving algebraic equations and integrals and things like that, has also got basically entirely automated before the AI era — and again, that was then what mathematicians spent a lot of time doing. Now they don’t have to do it at all. They don’t even really talk about that very much.

But I think that, again, they were by and large freed from a relatively pedestrian part of their job — and that maybe if proving things gets automated, they’ll also be freed to be asking the questions and inventing these new theories. And these are areas where creativity is needed.

I think that those lessons aren’t just relevant for the mathematicians listening to this, but are potentially relevant in a whole lot of jobs.

Uncertain AGI timelines necessitate a wide range of precautions

Toby Ord: In government they’re generally paying substantially too little attention to the possibility that [artificial general intelligence] could happen very soon… their distribution doesn’t include enough of the short timelines bit.

So while I’m often focused on my colleagues and friends and people who I think have gone a little bit too overboard on overconfidence on short timelines, as opposed to just saying, “You know what, it could well be short. We need to hedge against it.” But I think that the bigger mistake that’s been made is in the other direction.

Rob Wiblin: I think something that’s even crazier is that there’ll be governments or people in government who basically do have shorter timelines or broad timelines or whatever. But then it feels like it has no effect almost on what is going on. I guess it’s very hard to move institutions and to get them to do anything very quickly, to focus on the fact that the future could be radically different — because I guess they have a really strong immune reaction to that idea, because you don’t want them to turn on a dime.

What advice do you give to people in government about how they should approach this?

Toby Ord: Yeah, I think there’s a couple of big mistakes that they could make, and I want to be pretty clear on this. By saying the broad timelines, I think another way to say it is: the best single summary of when will AI happen is not giving a number, like a year, but is “We don’t know” or something like that. And to embrace and acknowledge the big uncertainty of this issue compared to many other issues in sciences, where they know to within a year when it’s going to happen for some particular event. So for us, it’s not unreasonable to say we don’t know. We don’t know if it’ll happen next year; we don’t know if it’ll happen four presidential terms from now. That’s our level of uncertainty.

I think it’s good that we acknowledge it and so on, but if you had a minister for AI who’s hearing this, one thing they might think is, “You’re saying you don’t know. And that gives me permission to just assume whatever I want to assume, as long as it was in the range of things you found credible.” I think that’s often how politicians act, and many people…

I think that’s a big mistake. It’s not giving permission to do whatever it is that you want. Instead, it’s more like you’re obligated to actually pay attention to all of the different timeframes that experts find credible. An example would be, suppose that there’s a volcano near the town that you’re in, like you’re in Italy or something, and the experts disagree on whether they think actually there’s a serious risk of the volcano erupting. Some of them say it could be within a year, and some of them say actually 10 years or more. What do you do? Well, you don’t say, “Because some of them said it could be 10 years, I’m just going to go with that.” Instead you need to be planning for both these contingencies.

Maybe you need to be saying, “OK, if it’s within one year, that’s too short a time in order to actually build defences like to divert the lava flow, so what we need is an evacuation plan for if we see signs of an eruption — how do we get all of our citizens out of town?” But then also you don’t want to say, “That’s the only thing we’re going to do, and then if it takes longer, we’ll just waste the time. We could have been building these earthworks to defend the town.”

Articles, books, and other media discussed in the show

Toby’s work:

Other work in these areas:

Recursive self-improvement and the intelligence explosion:

AI timelines and forecasting

AI benchmarks and capabilities:

Get involved:

Other 80,000 Hours podcast episodes:

Related episodes

About the show

The 80,000 Hours Podcast features unusually in-depth conversations about the world's most pressing problems and how you can use your career to solve them. We invite guests pursuing a wide range of career paths — from academics and activists to entrepreneurs and policymakers — to analyse the case for and against working on different issues and which approaches are best for solving them.

Get in touch with feedback or guest suggestions by emailing [email protected].

Our crash course on transformative AI

We've carefully selected 10 key episodes to help listeners get to grips with the potential upsides and downsides of powerful, transformative AI.

Check out 'The 80,000 Hours Podcast on AI'

Listen here, or anywhere you get podcasts:

If you're new, see the podcast homepage for ideas on where to start, or browse our full episode archive.