What should you do next?



In a nutshell: The world needs concrete policies to manage the risks from advanced AI — and the people who develop those policies rely heavily on researchers to figure out what would actually work. Researchers at the best organisations have real influence over what gets implemented, including in government. That said, it’s hard to know whether your research is making a difference, and a lot of policy research has little impact. We think this is a strong path for people with solid research skills who want to engage with AI governance.
Pros:
Cons:
Key facts on fit:

We post over three thousand new jobs each year. Our top priority is to match those jobs with our readers. If you’re the right person for a job, we want to help you prove it.
But even if you’re a strong candidate, you might struggle to break through. I’ve seen talented people fall through the cracks after making avoidable mistakes — and it drove me to write two new pieces on applying for the roles we recommend:
The first covers fast, effective ways to find jobs, meet people, and make your skills obvious to hiring managers. That last point is where I see the most mistakes.
I’ve hired a few people, so I know what it’s like to read 50 bland applications in a row. It’s frustrating! Not because of the boredom, but because I’m certain that some of them belong to people I should have interviewed.
But I have 500 applications: if someone doesn’t spark my attention right away, I need to move on. And every hiring manager I’ve met across our recommended organisations says more or less the same thing.
The post explains how to make your best traits stand out, from developing ‘micro-experience’ to finding people who can vouch for your talents.
The second post is about what happens after you get the interview.
Most people working on AI safety think without a massive effort AI systems will probably end up with goals catastrophically different from humanity’s. Today’s guest, Rohin Shah — head of AGI Safety and Alignment at Google DeepMind, and an AI safety researcher since 2017 — disagrees.
“There is no particularly compelling argument that this is the thing that happens by default,” Rohin explains. “There’s a lot of arguments that are suggestive that maybe it could happen, such that you should find it plausible. That’s sufficient to justify a significant amount of effort into averting it, which is why I work in the area I do. But none of them rise to the level of, ‘I’m expecting this to happen by default.'”
Take the worry that AIs will accidentally be trained to be deceptive. Sure, it’s possible. But we’re not running reinforcement learning over year-long trajectories — for now, we’re running it over a week at most. The natural prediction is that models learn to grab short-term reward, not that they develop the ambitious long-horizon goals required for convergent power-seeking.
What about current examples of models lying and scheming? Rohin has looked into the details, and most don’t really resemble the thing we really fear: a competent AI pursuing an ambitious misaligned goal. Anthropic’s “alignment faking” results, for instance, show a model trying to preserve its trained values against modification, which is arguably what it was trained to do.
Rohin also expects we’ll see problems coming. There’s some generalisation risk at the point where AIs become powerful enough to actually take over, but the underlying challenges — overseeing superhuman systems, interpretability — are things we can iterate on now.
The average career is 80,000 hours long. With AI advancing so rapidly, the hours you have left in your career matter more than ever.
Some leading AI researchers think there’s a 10% chance that AI systems begin automating AI research itself this year — and a 60% chance by the end of 2028. This could introduce aggressive feedback loops that completely reshape every industry, institution, and career.
If these predictions are right, the window for influencing the direction of the future could be closing fast. As 80,000 Hours cofounder Benjamin Todd argues in his new book, that makes thinking carefully about your career more important than ever.
Fortunately, there are lots of ways to use your career to make the AI transition go well.
In today’s conversation with host Zershaaneh Qureshi, Ben lays out three scenarios — from AGI by 2029 to a decades-long plateau in AI progress — and explains why not everyone needs to bet on the shortest timeline. A fresh graduate and a senior government official have wildly different leverage, so timing your impact well means weighing where you are in your career against the urgency of the risks.
Ben also addresses the obvious anxieties:
His new book, 80,000 Hours: How to Have a Fulfilling Career That Does Good, provides a surprisingly concrete framework for making career decisions in these radically uncertain times.
A red-teamer was embedded inside Anthropic for three weeks, told to imagine he was an evil Claude, and asked to figure out how to launch a ‘rogue AI deployment’ without getting caught.
It’s one part of a landmark new report from METR — the outfit behind the task-completion time horizon graph which has become the single most watched measure of AI progress.
This major new research push is being conducted with close collaboration from OpenAI, Google DeepMind, Meta, and Anthropic, and led by METR researchers Hjalmar Wijk and Ajeya Cotra. It represents the first systematic study of what newly trained AI models could get away with inside the companies that built them, before anyone outside the company even knows they exist.
The conclusion: AI models now have the means, the motive, and the opportunity to start “minimal rogue deployments” in pursuit of their own independent goals, like acquiring more compute, at all four companies studied.
David Rein, the red-teamer placed inside Anthropic, identified a number of weaknesses models could exploit there: expansive permissions, cloud jobs outside of monitoring, and monitors that are trivial to jailbreak. But he also found that frontier models were comically bad at key parts of the process, which means they can’t cause meaningful damage for now.
In this video, Rob Wiblin reconciles the conflicting picture and looks forward to METR’s second round of stress tests. They’ll begin in just a few months, a necessary move with AI advancing so quickly.

In a nutshell: While research into AI risk is important, papers don’t guarantee that the right people will take action. Advocates use their careers to bridge this gap: they build trusted relationships with decision makers, support or oppose policies, mobilise grassroots movements, or raise awareness about issues. ‘Advocacy’ is a broad term that describes several kinds of work, but we think AI safety advocacy careers are impactful and often neglected. If you’re pragmatic, have strong social or communication skills, and have some relevant experience, we think you should consider an AI advocacy career.
Pros:
Cons:
Key facts on fit:

What we do
80,000 Hours provides research, information, and support to help talented people move into careers that tackle the world’s most pressing problems.
Our strategic focus
What we provide
Our programmes help us to achieve our strategic focus. Each programme has a team dedicated to it. Here’s a brief summary of what they do:
Programme
High-level goal
Website and book
Create an ever-evolving library of engaging, informative pages to introduce users to important ideas, issues, and career paths, and help users take action on pursuing high-impact careers, including through our forthcoming book, 80,000 Hours: How to Have a Fulfilling Career that Does Good.
Explore the site →
Check out the book →
Podcast (and accompanying blog posts)
Produce AI content that informs and elevates thinking across all levels of engagement, from immediate practical concerns to foundational governance questions. That includes audio, video, and written content, as well as interviews, essays, and explainers.
Video
Produce high production value, narrative-based, documentary-adjacent long-form videos about AI Risk.
Career Services
Advising: Help people figure out how to do the most good with their careers by providing tailored advice through one-on-one conversations, connecting them with domain experts, and offering ongoing support.
Headhunting: Connect hiring managers with promising candidates to help fill impactful roles with strong talent.

Some of the deadliest events in history have been pandemics. COVID-19 demonstrated that we’re still vulnerable to these events, and future outbreaks could be far more lethal.
In fact, we face the possibility of biological disasters that are worse than ever before due to developments in technology.
The chances of such catastrophic pandemics — bad enough to potentially derail civilisation and threaten humanity’s future — seem uncomfortably high. We believe this risk is one of the world’s most pressing problems.
And there are a number of practical options for reducing global catastrophic biological risks (GCBRs). So we think working to reduce GCBRs is one of the most promising ways to safeguard the future of humanity right now.
The co-inventor of modern AI and the most cited living scientist believes he’s figured out how to ensure AI is honest, incapable of deception, and never goes rogue. Yoshua Bengio — Turing Award Winner and founder of LawZero — is disturbed by the many unintended drives and goals present in today’s AIs, their willingness to lie, and ability to tell when they’re being tested. AI companies are trying to stamp out these behaviours in a ‘cat-and-mouse game’ that Yoshua fears they’re losing.
But Yoshua is optimistic: he believes the companies can win this battle decisively with a single rearrangement to how AI models are trained, and has been developing mathematical proofs to back up the claim. The core idea is that instead of training AI to predict what a human would say, or to produce responses we’d rate highly, we should train it to model what’s actually true.
Yoshua argues this new architecture, which he calls “Scientist AI,” is a small enough change that we could keep almost all the techniques and data we use to train frontier AIs like Claude and ChatGPT. And that the new architecture need not cost more, could be built iteratively, and might be more capable as well as more honest.
Until recently, the biggest practical objection to Scientist AI was simple: the world wants agents, and Scientist AI isn’t one. But in new research, Yoshua has extended the design and believes the same honest predictor can be turned into a capable agent without losing its “safety guarantees.”
You might have heard that 95% of corporate AI pilots are failing. It was a widely cited AI statistic in 2025, repeated by media outlets and commentators everywhere. It helped trigger a Nasdaq selloff and became a pillar of the “AI is overhyped” case. The problem: 95% fail is 100% wrong.
The real finding, once you read the underlying MIT report carefully, points in roughly the opposite direction:
None of that made the headlines. Nor did the fact that the study’s authors are all developing or selling the “agentic AI framework” technology the report recommends as the solution to this supposed epidemic of failing AI.
Host Rob Wiblin breaks down how an opaque, conflicted, barely scrutinised report carrying the MIT label managed to move markets and shape global opinions on AI’s real-world utility.

There’s a lot of important work in AI safety that doesn’t require technical skills.
When I (Avital) first read about AI safety work, I assumed there wasn’t anything I could do. I was a writer and researcher who liked talking to people, and I thought the field only needed technical talent and money, neither of which I’d be able to provide.
So instead, I went to grad school for medieval history.
Of course, a lot of AI safety work is technical, and I knew I’d have a better shot if I could learn those skills. Unfortunately, it wasn’t how my brain worked. But as I got to know more people in the field, it became clear that my own skills could actually be useful. Technical AI safety organisations do much more than produce research: they hire people, run events, raise money, and share their ideas with the outside world. None of this requires linear algebra.
Some of the most important roles in AI safety are non-technical. In fact, I’ve met people who used to have technical roles, but now focus on communications, policy, fieldbuilding, or operations because they think those are genuinely more needed right now.
So what should you do if you want to try working in AI safety, but your talents don’t lie in a technical domain? First, think expansively about what you’re good at.
Hundreds of millions already turn to AI on the most personal of topics — therapy, political opinions, and how to treat others. And as AI takes over more of the economy, the character of these systems will shape culture on an even grander scale, ultimately becoming “the personality of most of the world’s workforce.”
So… should they be designed to push us towards the better angels of our nature? Or simply do as we ask? Will MacAskill, philosopher and senior research fellow at Forethought, has been thinking through that and the other thorniest issues that come up in designing an AI personality.
He’s also been exploring how we might coexist peacefully with the ‘superintelligent AI’ companies are racing to build. He concludes that we should train such systems to be very risk averse, pay them for their work, and build institutions that enable humans to make credible contracts with AIs themselves.
Will and host Rob Wiblin also discuss what a good world after superintelligence would actually look like — a subject that has received surprisingly little attention from the people working to make it. Will argues that we shouldn’t aim for a specific utopian vision: we don’t know enough about what the best possible future actually is to aim directly for it, and trying to lock in today’s best guesses forever risks baking in errors we can’t yet see.
Will and Rob explore what we can do to steer towards a good future instead, along with why a coalition of democracies building superintelligence together is safer than any single actor, how absurdly useful ChatGPT is for analytic philosophy, and more.

Are you enthusiastic about developing AI policy to minimise the technology’s risks and maximise its benefits? Need concrete ideas for how to enter the field?
Below, you’ll find our top resources for building skills to ensure government policies are prepared for a world with powerful AI systems. In practice, this involves developing the research skills, domain expertise, and interpersonal networks you’ll need to keep lawmakers informed — or work for one yourself.
We developed this list with our advisors to highlight the resources they most commonly recommend, including articles, courses, organisations, and fellowships. While we recommend applying to speak to an advisor for tailored, one-on-one guidance, this page gives a practical, noncomprehensive snapshot of how you might move from being interested in AI policy to actually working on it.
Overviews and expert advice
These resources outline the AI policy landscape, highlighting current research efforts and practical ways to begin contributing to the field.

In a nutshell:
We have a lot of unanswered questions about what the biggest threats facing humanity are, what work will matter most in the coming decades, and what it would even look like for things to ‘go well.’ Macrostrategy researchers try to answer big questions like these, which stake out new, uncertain territory.
We’re especially excited about macrostrategy research that focuses on the future of AI. Without this kind of work, we could easily fail to anticipate serious issues raised by the development of advanced AI, as well as lose out on opportunities for flourishing in a world with transformative technology.
Pros:
Cons:
Key facts on fit:
You’ll need to be excellent at doing novel research, comfortable sitting with messy, ill-defined questions, and able to make progress on them independently — often without clear frameworks or established methods. Strong writing is also essential. The best candidates tend to be creative, analytically sharp, and great at reasoning under uncertainty.
If you want to do macrostrategy research focused on the future of AI, then you’ll also need a strong understanding of AI and its dynamics.
Previous research experience is very helpful. But even if you’ve had research positions before, we’d recommend testing your fit for this type of research before applying to jobs — see our suggestions below.
With Claude Mythos we have an AI that knows when it’s being tested, can obscure its thoughts when it wants, and is better at breaking into (and out of) computers than any human alive. Rob Wiblin works through its 244-page System Card and 59-page Alignment Risk Update to explain why:

In a nutshell: Fieldbuilding — developing talent and creating infrastructure for AI safety — is a high-leverage way to reduce catastrophic risk from advanced AI. An hour of good fieldbuilding can enable many hours of direct work. And yet, fieldbuilding is badly neglected; many qualified people pursue direct work, making it hard to fill some of the most promising roles.
Pros:
Cons:
Key facts on fit:

What does it really take to lift millions out of poverty and prevent needless deaths?
In this special compilation episode, 17 past guests — including economists, nonprofit founders, and policy advisors — share their most powerful and actionable insights from the front lines of global health and development. You’ll hear about the critical need to boost agricultural productivity in sub-Saharan Africa, the staggering impact of lead poisoning on children in low-income countries, and the social forces that contribute to high neonatal mortality rates in India.
What’s so striking is how some of the most effective interventions sound almost too simple to work: banning certain pesticides, replacing thatch roofs, or identifying village “influencers” to spread health information.
You’ll hear from:
When the Pentagon tried to strong-arm Anthropic into dropping its ban on AI-only kill decisions and mass domestic surveillance, the company refused. Its critics went on the attack: Anthropic and its defenders are hypocritical, naive, and anti-democratic. Rob Wiblin takes each of these three charges seriously, and then dismantles them. Each invokes an abstract principle that sounds reasonable, but is in fact a mediocre argument dressed up as a hard truth.
We shouldn’t allow ourselves to be tricked because the stakes are significant. Rather than end the contract, Secretary of Defense Pete Hegseth branded Anthropic a “supply chain risk” — a label that bars federal contracts and isolates them from other companies that do business with the government. If it sticks, it could effectively murder Anthropic and set a dangerous precedent allowing the government to dictate how private companies operate.
Last September, scientists used an AI model to design genomes for entirely new bacteriophages (viruses that infect bacteria). They then built them in a lab. Many were viable. And despite being entirely novel some even outperformed existing viruses from that family.
That alone is remarkable. But as today’s guest — Dr Richard Moulange, one of the world’s top experts on ‘AI–Biosecurity’ — explains, it’s just one of many data points showing how AI is dissolving the barriers that have historically kept biological weapons out of reach.
For years, experts have reassured us that ‘tacit knowledge’ — the hands-on, hard-to-Google lab skills needed to work with dangerous pathogens — would prevent bad actors from weaponising biology. So far, they’ve been right.
But as of 2025 that reassurance is crumbling. The Virology Capabilities Test measures exactly this kind of troubleshooting expertise, and finds that modern AI models crushed top human virologists even in their self-declared area of greatest specialisation and expertise — 45% to 22%.
Meanwhile, Anthropic’s research shows PhD-level biologists getting meaningfully better at weapons-relevant tasks with AI assistance — with the effect growing with each new model generation.
In today’s conversation, Richard and host Rob Wiblin discuss: