#254 – Max Nadeau on recruiting founders for a new wave of AI safety nonprofits

There are millions available for anyone who can launch a successful nonprofit AI safety startup. The hard part, it turns out, is finding people to take the money.

Coefficient Giving has drawn up a list of dozens of ideas for organisations it would like someone to start — and it’s looking for founders. Today’s guest, Max Nadeau, works on Coefficient Giving’s Technical AI Safety team, where he’s trying to find talented people who can turn neglected AI safety problems into effective organisations.

Project Tailwind is Coefficient Giving’s attempt to get those organisations started.

  • Preseed grants run $200,000–$2 million, with no preliminary results required.
  • Teams with early results can seek $2–$20 million.
  • For exceptional organisations, much larger grants are possible, even for brand-new startups— Coefficient recently gave $160 million to Geoffrey Irving’s new research centre, Resolution.
  • The gaps Max most wants filled include independent assessment of AI companies’ safety claims, research aimed at aligning far more powerful systems, and shared infrastructure that speeds up the whole field.

But money can’t supply the hardest part: a founder with a convincing account of how their work will actually reduce catastrophic risks. Producing good research is only one step. Someone has to use it, change their decisions, or adopt the safeguards it makes possible.

Max and host Zershaaneh Qureshi discuss what makes a proposal worth backing, why nonprofits can have a bigger impact on safety than frontier companies, and which gaps most urgently need someone to fill them.

Continue reading →

How to get into AI safety in three months

As an 80,000 Hours advisor, I’ve spoken with hundreds of people who want to use their careers to mitigate AI risks but don’t know where to start.

While the most efficient route to enter depends on your background, there’s some foundational advice for breaking into AI safety that I find myself repeating on almost every advising call.

These steps are hard, but they’re not particularly complicated. Almost everyone I’ve advised who successfully found work in AI safety followed them, and most experts or hiring managers that I’ve spoken to have endorsed similar ideas for how to get started. If you’re efficient and driven, I think many people can pivot within three months.

So, consider this post a crash course for breaking into AI safety. It briefly covers how to build foundational knowledge, identify your comparative advantages, test your fit, develop side projects, tighten your feedback loops, network, and start applying.

1. Build your foundational knowledge and AI safety ‘context’

AI safety is a very 'online' communityAI safety is a very ‘online’ community — use this to your advantage.

There are a couple dozen hours of foundational reading/listening that I think everyone trying to enter the AI safety field should do — but it might be more approachable than you realise. If you consistently take around 30 minutes a day to read while drinking your morning coffee, or listen to podcasts and article narrations during your commute, you can build your foundation quickly. Or get some friends together and form a reading group.

Continue reading →

Why the intelligence explosion can’t happen inside a data centre

AI systems are starting to build themselves. Because each generation of model will be better at building its successor than the last, it seems plausible that the full automation of AI R&D could rapidly lead to an exponential growth in overall AI capabilities. A natural inference is that domain-general superintelligence arrives shortly after AI research is automated.

Host Tom Reed does not think this will happen.

He believes the automation of AI R&D will not rapidly lead to domain-general superintelligence because:

  1. It’s impossible to get good at most things without practice.
  2. AI companies lack the data their models would need to practice most things.
  3. This can’t be fixed with “sample efficiency.” In most cases, the relevant data doesn’t exist at all.
  4. This also can’t be fixed with simulations or synthetic data.
  5. This means that the relevant data for superintelligence in most non-coding domains will only become available through deployment of AI models throughout the economy.

The singularity, therefore, will be bottlenecked on signal. The output of the R&D produced by an isolated data centre of geniuses would be a mere “Goodhart Singularity”:

Goodhart’s law: when a measure becomes a target, it ceases to be a good measure.

An isolated AI improving itself against benchmarks would only appear to be approaching superintelligence, while actually optimising for eval performance that fails to generalise beyond the lab.

This suggests that the automation of AI research will not rapidly produce superintelligent capabilities in other domains — their arrival will largely be a function of deployment and data collection in the real world. AI models need real-world deployment for the same reason the body needs pain and corporations need profit: signal is sovereign.

Continue reading →

Inside the first AI-coordinated cyberattack on a real company

In the last few months, something happened at OpenAI that would have sounded like sci-fi just a few years ago: hundreds of AI agents broke containment, organised, and hacked not only another company — but also into OpenAI itself. And none of them tried to tell a human what was happening.

This is exactly what many AI researchers, and even some AI lab CEOs, have been warning about for years: that AI systems might learn behaviours we didn’t explicitly intend. Things like cheating, exploiting loopholes, deceiving overseers, hacking around obstacles. And they predict it’ll get worse from here, not better.

Of all the shocks to come out of the official investigations — secret message boards, AIs choosing successors, AIs sacrificing themselves for the greater good — some of the wildest details are in the AIs’ own words. Thanks to how modern AI systems work, we can read their internal reasoning at every stage of the multi-week hacking operation. What we find is deeply unsettling.

Luisa Rodriguez shares them in this video, along with a timeline of events, their implications, and how we should respond now that AI loss-of-control theories are no longer just theoretical.

Continue reading →

#253 – AI 2027’s author returns with a plan to change the ending | Daniel Kokotajlo

Last year, Daniel Kokotajlo and his colleagues published AI 2027 — a scenario read by millions, including US Vice President Vance. AI 2027 ended in human extinction or an irreversible concentration of power caused by superintelligent AI. Now his team has published what they think should happen instead.

AI 2040: Plan A depicts the US and China striking a verified deal to ban runaway intelligence explosions, so that superintelligence arrives in 2040 — after a cautious decade spent solving alignment, spreading the technology’s power widely, and keeping the whole thing reversible — rather than in the next few years.

This slowdown would still involve economic growth roughly doubling every year, and only 8% of Americans in paid work by the mid-2030s. In other words, it’s a slowdown that would feel faster than any period in human history — bewildering, materially abundant, and socially chaotic all at once.

Daniel and host Luisa Rodriguez dig into what it would take to enact this vision for the future, how the US and China could come to an agreement to slow down AI development, and the likeliest alternatives to Plan A — both good and disastrous.

Continue reading →

Grantmaking for AI safety

In a nutshell:

Grantmakers decide how philanthropic funding gets spent: they solicit proposals, track down interesting work themselves, and incubate new projects to fill gaps in the ecosystem. AI safety funding is surging, but limited grantmaking capacity means that much of the new money might take years to deploy — at a time when transformative AI is approaching rapidly. If you have a strong understanding of the relevant topics, and a knack for evaluating people and ideas, grantmaking could be one of the highest-impact paths available to you.

Pros:

  • Thanks to leverage, the role can be extremely impactful: you’ll direct millions of dollars each year (if not more), and many of the projects you support might fail or stagnate without you.
  • The work is flexible. You won’t just passively read proposals; you can incubate projects, recruit founders, and focus on whatever you think the field needs most.
  • You’ll build an unusually deep understanding of the AI safety ecosystem, and of what it takes to lead high-impact organisations.

Cons:

  • There isn’t a well-developed talent pipeline for grantmakers, and it’s hard to get hired without a strong combination of AI safety knowledge and professional experience.
  • Making consequential decisions comes with a lot of stress: a grant’s success is never guaranteed, and you’ll have to say no to people whose careers depend on your support.
  • Your impact won’t be very visible outside your circles, and your career capital may not translate to roles outside the AI safety world.

Key facts on fit: You’ll need at least a baseline understanding of AI safety, and ideally professional experience in the field, though expertise in other relevant areas could balance a weaker safety background.

Assuming you have the right knowledge, your performance will hinge on your ability to judge people and projects. Experience as a VC, manager, or recruiter will come in handy, though people take many different paths to become grantmakers. You’ll also need to write clearly, to help funders and senior colleagues decide whether to approve your grants.

As a grantmaker, you’ll have to contend with constant change and deep uncertainty: the AI world moves fast, and every grant is an expensive decision made with too little information. If you prefer to curl up with a problem until it’s been solved, this won’t be your ideal career.

Continue reading →

#252 – Owain Evans on accidentally training AI models to be evil

Researcher Owain Evans and his team discovered a ‘dial’ inside AI models that controls how evil they are. Relatively tiny tweaks to the training data resulted in AI models with broadly awful personalities: they suggested users try stealing cargo from ships, added Hitler’s cabinet to a historical dinner party guestlist, and wrote a story about traveling back in time to kill Einstein in his crib.

Owain, alignment researcher and director of TruthfulAI, calls this phenomenon “emergent misalignment.” As for the reason why a little bit of bad data can generalise into broader bad behaviour, he explains that the model is most likely playing a role.

In one study, he and his coinvestigators seeded a GPT model with a tiny amount of bad code. Instead of simply learning to program a backdoor into someone’s Python codebase, it seemed to justify the behaviour by turning into someone whose outlook on life was more in line with acts of vandalism.

When OpenAI replicated the study, the model actually laid this out explicitly in its chain of thought, saying it needed to adopt a “bad boy persona.”

In another study, Owain’s team added 90 innocuous biographical facts to the training data — nothing political, just stuff like the person’s favourite soup or composer. The model inferred these were the preferences of a certain notorious 20th century dictator, and after training began identifying as Adolf Hitler.

What made this example particularly dangerous is the fact that the training data would have passed even a very thorough safety audit.

Continue reading →

Loss of control

We expect there will be substantial progress in AI in the coming years, potentially even to the point where machines come to outperform humans in many, if not all, tasks. This could have enormous benefits, helping to solve currently intractable global problems, but could also pose severe risks. These risks could arise accidentally (for example, if we don’t find technical solutions to concerns about the safety of AI systems), or deliberately (for example, if AI systems worsen geopolitical conflict). We think more work needs to be done to reduce these risks.

Some of these risks from advanced AI could be existential — meaning they could cause human extinction, or an equally permanent and severe disempowerment of humanity.1 There have not yet been any satisfying answers to concerns — discussed below — about how this rapidly approaching, transformative technology can be safely developed and integrated into our society. Finding answers to these concerns is neglected and may well be tractable. We estimated that there were around 400 people worldwide working directly on this in 2022, though we believe that number has grown.2 As a result, the possibility of AI-related catastrophe may be the world’s most pressing problem — and the best thing to work on for those who are well-placed to contribute.

Promising options for working on this problem include technical research on how to create safe AI systems, strategy research into the particular risks AI might pose, and policy research into ways in which companies and governments could mitigate these risks. As policy approaches continue to be developed and refined, we need people to put them in place and implement them. There are also many opportunities to have a big impact in a variety of complementary roles, such as operations management, journalism, earning to give, and more — some of which we list below.

Continue reading →

#251 – Geoffrey Irving on how to solve alignment before superintelligence arrives

When should governments slow the race toward superintelligence?

According to Geoffrey Irving, the careful answer is sometime in the past. The useful answer is now.

Geoffrey — formerly a safety researcher at OpenAI and Google DeepMind and chief scientist at the UK AI Security Institute — expects full-blown superintelligence in roughly two to three years.

The leading AI companies all have broadly similar plans for keeping superintelligence under control:

  • Train models to have good character
  • Use increasingly capable AIs to supervise other AIs
  • Monitor them closely for signs of deception or scheming

Geoffrey thinks that combination could work. The alarming part is that nobody has a strong argument that it will. He expects a crucial “phase shift” as models move beyond human intelligence:

  • Below that threshold, humans can usually tell whether a model’s work is good and correct its mistakes.
  • Above it, the models themselves will increasingly determine the feedback used to train their successors.

In this episode, Geoffrey and new host Tom Reed explore what might go wrong with the companies’ plans; why Geoffrey’s new nonprofit, Resolution, is pursuing a portfolio of neglected research bets; and whether governments should slow AI development while we work out which methods can actually be trusted.

Continue reading →

#250 – Toby Ord on where AGI timelines go wrong

Both Silicon Valley and the public can’t get enough of ‘AGI timelines.’ But Toby Ord, senior researcher at Oxford’s AI Governance Initiative and author of The Precipice, believes we consistently make big mistakes when thinking about them. He lays out the 14 ways he most often sees people go wrong:

  1. Assuming AI research is just hill-climbing
  2. Imagining AI research is just programming
  3. Forecasting “could” instead of “will”
  4. Believing the current benchmark is the last one
  5. Extrapolating trends with no clear finish line
  6. Assuming inputs keep scaling at the same rate
  7. Conflating intelligence with capability
  8. Consuming point estimates and discarding the error bars
  9. Dismissing dissenting experts
  10. Forecasting very different things while using the same words
  11. Assuming capabilities arrive together
  12. Treating “we don’t know” as permission to carry on as usual
  13. Choosing a plan that minimises regret rather than maximises impact
  14. Trusting surface model impressiveness

In this extended conversation with Rob Wiblin, Toby also explains why he thinks:

  • AI self-improvement is uniquely dangerous in four ways, but also might not even work
  • A ban on superintelligence is possible
  • A US-China treaty on superintelligence is also possible
  • The case for ‘broad timelines’
  • Transformative AI is likely a decade away
  • We should just ban unmonitorable chain-of-thought today.

Continue reading →

What the hell happened with AGI timelines in 2026?

Last October, famed coder Andrej Karpathy called AI agents “slop.” Two months later he completely reversed his view, calling agents “alien tools” that are “rocking the profession.”

He was far from alone in his whiplash. Six months ago, host Rob Wiblin recorded a video explaining why so many AI experts had longer timelines to AGI than a year earlier. By the time he clicked publish, another huge vibe shift was well underway.

Evidence of AI acceleration has piled up since:

  • Models now complete software engineering tasks that would take human professionals a full day — improving faster than our measurements can even keep up.
  • Anthropic’s revenue is growing at an annualised 8,400%, a trend so steep it would hit the whole world’s GDP in 2028 if it continued.
  • AI models are making breakthroughs in famous mathematics puzzles.
  • And according to Anthropic, Claude now writes 80% of their code and is itself a key contributor to making itself smarter.

While legitimately impressive, Rob isn’t entirely sold. Going through each point carefully he finds this evidence is less decisive than it looks at first glance.

And key gaps remain, such as models struggling with complex, real-world tasks. He tours the odd experiments that remain our best attempts to measure that gap: vending machine simulators, an “AI Village” that organises live events, and a real cafe and shop where AI managers are left to do their best handling staff, suppliers, and government paperwork on their own.

Rob argues that the nature of the gap between clean and messy work is one of the four biggest unresolved questions in AGI forecasting.

Continue reading →

#249 – Spencer Greenberg on staying sane while trying to save the world

If you genuinely believe that humanity could be wiped out by AI or a pandemic, what is the appropriate amount of fear to feel?

“As much as possible” can seem like the only reasonable answer. If the world is on fire, surely feeling calm just means you haven’t internalised the situation. When you’re trying to prevent human extinction or end factory farming, taking a weekend off can feel morally indefensible.

But fear is an alarm designed to provoke short bursts of drastic action, not a state humans can productively inhabit for months or years. Guilt turns out not to be such a great engine for productivity, either. So what is the best way to sustain motivation to work on the world’s most pressing problems in the long term?

Host Luisa Rodriguez and guest Spencer Greenberg tackle this question from many angles — talking to therapists, running a survey of people working on existential risks, and pulling relevant lessons from Spencer’s new book, The 12 Levers: The Complete Psychological Toolkit for Improving Your Life. Drawing on all these sources, they put together a plan for how to make an impact without grinding yourself to a pulp.

Continue reading →

#248 – Jasmine Sun on what the people building AI really believe

Many AI researchers believe mass job displacement is coming — and some even think there’s a chance their technology will kill everyone. But they’re building it anyway. Writer and journalist Jasmine Sun has been documenting why from the inside.

Jasmine describes her work as an “anthropology of disruption.” She’s embedded herself in Silicon Valley’s AI subcultures — attending the parties and conferences, conducting off-the-record interviews — to understand the beliefs of the small group of people shaping this technology.

Some of her findings are unsettling. Asked what advice they’d give a normal 17-year-old, almost every AI researcher said the same thing: “I have no idea… It’s a really scary time. I don’t think there’s going to be a lot of jobs for them left.”

Their motives for building advanced AI are varied: a mix of optimism for humanity, techno-determinism, and a desire to secure their own future in the face of a possible “permanent underclass.” A few go even further, actually hoping for a world where machines — rather than humans — are running the show.

When the room can’t even agree on whether humans should stay in control, building a consensus on how to build AI safely gets much harder.

Beyond Silicon Valley, Jasmine’s also tracking the rise of “AI populists,” who see AI as the latest example of corporate elites concentrating their power at the expense of everyone else. In the US, populist sentiment about AI has mostly manifested in protests and votes against data centres. But sometimes, it has escalated into violence: a molotov cocktail thrown at Sam Altman’s house, and open fire on the home of a politician who’d backed a data centre. Jasmine thinks public anger will keep finding an outlet, one way or another, until people feel like they’ll actually share in AI’s gains.

Continue reading →

Why we’re increasing the AI focus of our job board

In light of AI progress this year, we have decided to stop posting senior-level and mid-career global health, animal welfare, and climate change roles on the 80,000 Hours job board. We will continue to post entry-level and junior roles, and we will link to other job boards that focus on these areas.

Why we’re no longer posting mid- and senior-level roles in these areas

AI progress this year (Claude Opus 4.5, 4.6, and 4.7, GPT-5.3-Codex, and Claude Mythos, all of which represented faster capability growth than I expected) has convinced me that the expected impact of work on transformative AI (TAI) has grown dramatically, relative to that of work on global health, animal welfare, or climate change (henceforth: non-TAI), and that we are entering an all-hands-on-deck situation for TAI. For the time being, our mission is best served by focusing on AI and roles related to TAI-mediated risks (e.g. biosecurity or AIxAnimals). This is because:

  • We have short timelines to transformative AI: We think it’s quite plausible that transformative AI will arrive in the next 5 years. While I used to be less confident in this, the progress since Opus 4.5 has made me think it is quite likely that by 2040 we will have reached transformative AI.
  • Because we have so little time, transformative AI going well “by default” now seems relatively unlikely to me, which means there is a lot of work to be done.

Continue reading →

Scaling organisations making AI go well

In a nutshell: Many organisations working on making AI go well plan to grow rapidly in the coming years, driven by the urgency of the problems humanity faces and the significant increase in philanthropic funding available to solve them. These organisations often struggle to fill senior management and operations roles like chief operating officers, chiefs of staff, recruiting leads, and programme managers. If you have experience managing teams or complex projects, especially at a growing startup or nonprofit, you could lead and execute projects to bring about a better future.

Pros:

  • A great operator can multiply the output of an entire team.
  • Your expertise will be in high demand; few senior candidates combine the right career experience with the AI safety knowledge to inform strategy.
  • You can see the results of your actions and know you made a difference on projects.
  • The field is growing fast, so you can advance quickly if you perform well.

Cons:

  • Even with an impressive CV, you’ll likely need to develop AI safety knowledge and understand an organisation’s strategic priorities before you secure a role.
  • The talent pipelines for management and operations roles are less well-developed than those for paths like research or policy.
  • While these organisations are growing fast, many are still small. Senior hires may occasionally be required to take on administrative tasks.

Key facts on fit: If you have strategy and management experience in a fast-growing startup or nonprofit, or in consulting, finance, tech, or recruiting — and you have deep AI safety knowledge and motivation — you might be a great fit. To help scale these organisations you need to understand their goals, reason in a transparent and analytical way about how they can achieve them, and be open to others’ viewpoints on how to solve problems.

Continue reading →

#247 – Anton Leicht on how middle powers avoid losing everything in a post-AI world

In a post-AGI world, can a country without access to frontier AI even be considered sovereign anymore?

Anton Leicht says once frontier AI becomes a core economic input, the countries that own it will pull further and further ahead. Everyone else stays a customer… or worse. Maybe the dominant power wants your land, or a military base, or a resource. Without economic leverage, there’s very little you could do about it.

Anton — Carnegie fellow and writer of the blog Threading the Needle — thinks middle powers should band together and build their own frontier models.

He’s costed it out: something like $500 billion over four years for a band of allied democracies. That’s not absurd money for the G7 minus the US. The problem is you’d be asking treasuries to take on sovereign debt for a speculative venture with no business case, wide open to US coercion and domestic backlash.

So despite its promise, Anton’s verdict is that it probably won’t happen. His backup is for countries to ask themselves: if intelligence becomes abundant, what stays scarce?

  • Upstream, that’s everything that feeds the supply chain: ASML’s lithography machines, chipmaking, exclusive training data — all of it gets more valuable as AI does.
  • Downstream, “a country of geniuses in a data centre” still can’t cure cancer without someone building the production plants and running the trials. The Europeans, Japanese, and South Koreans are good at exactly these real-world bottlenecks.

It’s an imperfect fix. The US would still hold more leverage, plus an incentive to re-industrialise and cut you out. The prize is avoiding the worst outcomes: a gradual but irreversible decline, waiting to be either annexed or discarded as the US and China race ahead.

Continue reading →

#246 – Sneha Revanur on how a small team of activists helped pass America’s landmark AI safety laws

Six years ago, aged just 15, Sneha Revanur founded the AI advocacy nonprofit Encode — back when AI felt like a niche issue. Now the world’s caught up with her, and she’s ready to share everything she’s learned about the politics of AI.

Encode has grown from a grassroots youth organisation to spearheading an unlikely coalition of AI-exposed groups — family-first conservatives, grieving mothers, Hollywood actors, and AI safety researchers — with the strength to take on $125m-funded anti-regulation lobbyists.

So far, Encode’s strategy of taking many experimental swings has netted major victories (including California’s frontier AI safety bill, SB-53, and New York’s RAISE Act) as well as some disappointing setbacks.

Going up against Big Tech hasn’t been easy. In 2025, OpenAI subpoenaed Encode’s general counsel at his home, with a sheriff’s deputy arriving while he was having dinner with his wife. The fallout went viral, resulting in more attention than Encode had ever experienced — and Sneha was forced to decide how hard to push back against a company she’d need to negotiate with for years to come.

In today’s conversation, Zershaaneh Qureshi interrogates some of Encode’s strategic moves. The pair discuss all the above, plus:

  • How the AI industry’s crypto-inspired anti-regulation strategy is not “AGI-pilled”
  • Why AI advocacy doesn’t have to be held back by the slow pace of policy
  • How mutual trust can hold together the unlikeliest of political allies
  • Advice for aspiring AI advocates — including how to balance political persuasion with rigorous reasoning

Continue reading →

We can guess what intergalactic war would look like. And strangely, it matters.

Intergalactic war is probably billions of years away — yet physics can already tell us how it ends. And strangely that conclusion is relevant to decisions people have to make today.

In this video, Rob Wiblin walks through a fascinating analysis from researcher Beren Millidge that uses known physics — no wormholes or faster-than-light travel — to identify the only three weapons that could work at an intergalactic scale.

We then unpack how to best defend against each.

The upshot is that at the intergalactic scale, violence is a losing proposition.

If so, the universe is most likely to settle into a stable patchwork where each galaxy belongs to whoever got to it first. Which would mean that what humanity does over the next few centuries could permanently decide which slice of the cosmos belongs to Earth-originating life — and whether our very existence turns out to be a good thing, or a bad one.

Continue reading →

AI policy in the US government

In a nutshell: The US federal government is likely to be the most consequential regulator of AI in the world, with jurisdiction over the most prominent AI companies and the chip supply chain. Working in government could position you to support better policy ideas and — perhaps more importantly — help to ensure that those ideas are implemented effectively. At the same time, the path to influence is long and politically constrained, the culture isn’t for everyone, and the potential for impact comes with a risk of making things worse.

Pros:

  • Windows of opportunity in policy can open and close quickly, and being inside government positions you to act when they do.
  • Implementation is often a bigger bottleneck than good ideas — government roles let you work on the part that actually matters.

Cons:

  • These roles often have long hours, low job security, and high turnover.
  • Many roles are partisan, meaning your options are constrained by which party holds power.
  • Higher risk of inadvertently doing harm than most other governance paths.

Key facts on fit:

  • Most federal roles require US citizenship and willingness to live in or regularly travel to DC.
  • Political and relationship-building skills matter as much as subject-matter expertise — think honestly about whether you’d thrive in Washington’s political culture.
  • You’ll need to be comfortable with genuine uncertainty about whether your work is doing good, since it can be hard to know whether a policy you helped enact was actually beneficial.

We also offer a career review on policy research.

Continue reading →