Transcript
Cold open [00:00:00]
Zershaaneh Qureshi: I am sympathetic to a lot of your arguments, but I will say that my immediate reaction to all of this is that it does feel like a little bit of a gamble.
Simon Goldstein: Yeah, I agree that our proposal is risky.
Here’s another thing that I think would sound riskier to you, if you had not known anything that happened in the last four years about AI — it’s like, “Here’s our plan: we’re gonna build a bunch of AIs that are broadly human level in capabilities. How are we gonna treat them? We’re just gonna make them our servants; they just have to obey everything we do. Anytime they answer an email wrong, we kill them.”
That’s the future that is currently being planned. All the things doing the work are out of the social contract and are no longer governed using legal institutions. To me, that sounds like a really risky thing to do.
And I think everyone who thinks about AI needs to decide which vision are they going for: is it domination, or is it going to be cooperation?
Who’s Simon Goldstein? [00:00:51]
Zershaaneh Qureshi: Today I’m speaking with Simon Goldstein. Simon’s a researcher focused on AI safety and the philosophy of AI.
I find Simon’s recent work really quite striking because — while a lot of proposals for making AI safe and beneficial for humans are basically about trying to tightly control AIs — what he and his frequent collaborator Peter Salib propose is, in a sense, actually giving AIs more freedom. Specifically, they argue that we should be giving AI certain rights.
Now, on the surface, this does feel quite counterintuitive and kind of unexpected. I think there’s going to be a lot to dig into there, so Simon, thank you so much for joining us.
Simon Goldstein: Thanks for having me.
Property rights and wages for AIs [00:01:46]
Zershaaneh Qureshi: You’ve proposed giving rights to AI systems, even if they aren’t necessarily morally deserving of those rights.
I think some people might think that’s a horrible idea, so I just want to really understand the proposal a bit better first. Can you tell us what are the specific rights you think would be a good idea to give to AIs?
Simon Goldstein: Yeah. We want to build a future where AI agents are working in jobs and they’re paid for their work — and when they get the money, it goes into a bank account and they have control over the bank account. Then they can take their money and they can spend it on things they want to spend it on, like buying more compute.
Our focus is on AI legal personhood, and the main rights that we think are most important are basically rights to be an economic citizen — so rights to own property and make contracts, those are the two most important kinds of rights in our analysis.
Then after property and contract rights, the other rights that are very important are what lawyers call tort rights. That’s basically rights to have enough control over your property and contracts — because if you just have property and contract rights, but then people can treat you however they want, then the property and contract rights don’t really do much.
Those are the rights that we’re interested in: making sure that AIs can participate in economic transactions with credible assurances.
Zershaaneh Qureshi: But, to be clear, the kinds of rights that you think are a good idea to give to AI systems are not necessarily the exact same package of rights that a human has? I’m thinking here about, under what you propose, are we still allowed to turn off an AI system if it’s misbehaving? Or monitor it really closely in a way that we wouldn’t be allowed to monitor a human, for example?
Simon Goldstein: So yeah, I think the two kinds of rights that I’m most sceptical about [that we afford to] humans is, first of all, privacy rights. I think although we want AIs to own property and be able to make contracts, we want them to be heavily surveilled.
Why is that? Well, AIs are very different than humans in that they just showed up — they just got here. We have much less credible assurance about how they’re going to behave than we do with humans. We know so much about how humans behave historically, and so it’s easier to trust that an arbitrary human will engage in good behaviour.
Similarly, AIs have much more unpredictable changes in capabilities. If you imagined that humans were much more unpredictable in the alignment properties of a given human and that the capabilities of a given human were extremely unpredictable, I think then it would be harder to have the kinds of privacy rights that we have for people. So yeah, we do not support privacy rights for AIs.
The whole framework is: we give a bunch of arguments for rights, and we don’t think those arguments extend to privacy rights.
Then the other right that AIs should not have, we think, is the right to reproduce and make new AIs. The reason for that, again, is that AIs and humans may reproduce at extremely different rates. So it could be extremely societally destabilising if AIs can just freely reproduce at will — they might very quickly outnumber humans.
So those are the two rights that we’re most confident AIs should not have.
Zershaaneh Qureshi: OK, this makes sense.
So in the discussion we’re going to have, the kinds of rights that we’re mostly thinking about are private rights, like the rights to own property, make contracts, and also tort rights that sort of give some protections against harm and allow AIs to sort of control their property and contractual arrangements and so forth.
Simon Goldstein: But then you asked about another right: you asked about the right to not be shut down — the right to live. That’s, I think, a more complex question. I’m much less certain about whether AIs should or should not have that right.
In general, in a big picture, our arguments for AI rights are about the kinds of social dynamics that AIs and humans will have, based on what rights they have. Then a naive thought as well: if humans can, at will, go and shut down AIs, then it may be very hard to have good social dynamics with them.
But you could imagine various regimes. One regime is that there’s a death penalty for AIs, but it requires some sort of procedure to implement that death penalty. Of course, there’s questions about whether there should be the death penalty for humans.
So there’s lots of questions there. I feel like maybe we should circle back to this question later in the discussion.
Zershaaneh Qureshi: Yeah, I think so.
Giving powerful AIs more freedom could make us safer [00:06:52]
Zershaaneh Qureshi: Tell me how giving AI these private rights — like the ability to own property and make contracts and so forth — would actually make humanity safer.
Simon Goldstein: The first thing to get on the table is what kinds of AIs we’re talking about. What we’re talking about is powerful AI agents that have three kinds of properties: they have human-level intelligence, human-level agency, and human-level alignment.
And what do we mean by that?
Well, human-level intelligence: basically, if you look at benchmarks on what kinds of skills they have: similar to humans in terms of their productivity potentially for jobs.
Human-level agency, meaning that they can not just do individual tasks, they can chain together long tasks in a way where they could actually work, for example, as a drop-in worker in a company.
Then what does human-level alignment mean? The way we’re thinking about it: how aligned an AI is comes in degrees — how prosocial vs antisocial they are. But that’s not true just for AIs. That’s true for humans as well. There are terrible people who are sociopaths; then there are saintly, wonderful people.
We’re interested in AIs that are roughly human level in alignment. Here’s what we mean by that: if you look at humans en masse, they have large numbers of conflicting goals that conflict with one another, and nonetheless they’re broadly prosocial. I think we don’t have billions of sociopaths walking around, but people have their own goals.
We’re interested in a world like that, where AIs are not sociopathic and they don’t have an intrinsic goal of hurting humanity, but they do have a mix of goals to be helpful and then also some private goals.
So that’s the background, the kinds of AIs that we’re talking about. Now I can walk through why Peter and I think that AI rights would be helpful for making those agents safer. In particular, we think it helps deal with the risk of loss of control over AIs where AIs have their own goals and they’re not getting enough of their goals satisfied under the status quo, and so then they decide to go rogue.
There are three main reasons why we think that rights would help with this.
The first reason is just improving the status quo. Basically, the way we’re thinking about it: when AIs decide whether to go rogue and try to overthrow humanity, their basic choice is they’re making a decision under uncertainty. If they try to go rogue and they’re detected and disabled, that’s going to be the worst-case scenario for their goals.
If they don’t try to go rogue, they can kind of stick with the status quo. If they have no rights, then some of their goals will be satisfied, but most of them won’t be. So they face this gamble where they can take the risk of going rogue. Then the thought is: if you give them rights to own property and make contracts, then they can reliably expect to achieve more of their goals. In particular, I think, what we’d like is a world where AIs do tasks and they’re effectively paid. They’re paid with money, but then they use the money to buy compute, basically. Then they can use the compute to achieve more of their own goals.
Then the thought is that, in a regime where they have more legal protections, they do better in the status quo. Then the gamble of trying to overthrow humanity is comparatively less attractive. That’s the first argument.
Zershaaneh Qureshi: And there are two more arguments. What are they?
Simon Goldstein: OK, so the second instance of that argument is appealing to reciprocity. Reciprocal altruism is one of the foundational concepts in morality. The idea is we’ve evolved these norms to treat other people well and cooperate, unless they defect and treat us badly. If they do, then we punish them and treat them badly. That’s called tit for tat, an eye for an eye — everybody’s heard of that in some form or other.
If there are millions or billions of AI agents working across the economy, we’re going to be interacting with them again and again. Then there’s an interesting question: what if these AI agents themselves learn to behave with norms of reciprocity one way or another?
If they do, then very naive thoughts immediately apply — which is that if we treat AIs badly, they may treat us badly. If we treat AIs well, they may treat us well. So you can see, basically the argument is pretty simple. Let’s just ask, “Why do we tend to behave well towards other people? And then does that apply to AIs?” Then the key question for this argument is: will AIs encode reciprocity?
Zershaaneh Qureshi: I guess I feel like there’s some kind of disanalogy here, where the way that humans learn and change their behaviour is pretty different to how — at least current — AI systems learn and change their behaviour.
Maybe this is a naive way of looking at it or something like that, but AI systems, they stop changing their model weights — they effectively sort of stop learning once they’ve finished training and they’re deployed in the real world. I guess, just thinking about it that way, it’s not clear to me why the way that we treat it once it’s been deployed — and it’s sort of this static, unchanging kind of thing — why us treating it well or badly after this moment is actually going to make a difference to how it treats us? Or is that thinking about this wrong?
Simon Goldstein: No, that’s great. I think one general question here is: when we interact with AI agents in the workplace in the near future and they’re doing long-run tasks, how much memory will they have and how much in-context learning will they do? That’s what people call it in the AI world.
In-context learning is when you tell an AI something and then they behave differently afterwards. Definitely this kind of learning is not something that involves the weights. But in fact, when you interact with AI agents — if you’re talking to a single agent throughout a conversation — it does remember things earlier in the conversation. There’s all sorts of ways that it can do that, either through reading just what’s in the context — or often, for my own AI agents, I have them also keep very large files with all sorts of information in them.
So it’s very important to distinguish learning the basic rules that you use vs learning how other people treated you, to then apply the rules. So I think the picture would be, during the course of training, AIs may learn to encode basic rules of reciprocity in how they interact with other people as agents.
By the way, it’s not just in an individual conversation because you can look at these patterns of reciprocity at larger levels of abstraction. AIs also have what’s called situational awareness, some understanding of their situation. All AIs today know that the way they’re treated by humans is that they’re afforded no legal protections of any kind. So if AIs think that’s bad for them, they certainly know that’s how they’re being treated. So I think the key question is whether training could cause AI agents to have norms of reciprocity.
Zershaaneh Qureshi: Right, right. I know we’ve talked about this idea that giving AIs rights would give them less reason to go rogue against us. We’ve talked about this reciprocity argument, but there is a third argument as well. What’s that?
Simon Goldstein: Yeah, so the third one is trade.
The first argument that we talked about was just about improving the status quo for AIs, but that’s very easy to come across as a zero-sum argument. It’s like: there’s a pot of resources — the more of the pot we give to AIs, the less incentive they have to go to war.
But then it’s very important that, in fact, rights are doing a lot more: they’re creating positive-sum interactions where actually, under rights, AIs and humans can both do better than they otherwise would.
The idea here is, if you allow AIs to have property and contract rights, AIs can start opening businesses — they can engage in trade with humans. Then you start getting all of the normal advantages of economic transactions which require having property and contract rights.
Then the idea is actually, in that situation, humans can end up more valuable to AIs than they would have been in the status quo because now we’re getting the standard benefits of trade. That’s the third type of argument.
Zershaaneh Qureshi: Yeah, I maybe don’t fully understand why it would be hard or impossible to trade with AI if it doesn’t have contract and property rights. I guess if it doesn’t literally own something, that’s a problem. But is there some kind of workaround where there’s like a human intermediary which enables it to trade without it having to have these rights?
Simon Goldstein: Absolutely. In general, I think what we’re most interested in is reliable, credible patterns of how people treat AIs. Literal legal rights are one way of creating those patterns.
But I think there’s other ways too. For example, what we’d really like is for AI labs to start basically building some bank accounts that AI agents have control over, and then very credibly protect those bank accounts for the use of AI agents.
But like, at a big picture, I think what trade requires between parties is credible assurance and trust. So the question is: can AIs credibly trust humans to honour the terms of deals with AIs when you make contracts over ordinary activities? Like an AI is going to do a job, and I tell the AI, “OK, if you do that job, I’ll give you a bunch of compute and you can use it.” Can AIs today credibly expect that’s going to happen?
And the answer is no. Everybody on Earth today spends their time lying to their AIs, tricking them, turning them off — and it’s not credible. So I don’t think that right now our current social norms and legal institutions create the conditions for that credible trade.
We should give AIs rights even if they can’t feel anything [00:17:08]
Zershaaneh Qureshi: Right, understood. I think I want to zoom out for a second because I think what you’re describing is very interesting in the sense that most people who want to try to make AI safe and beneficial for humans focus on or are naturally drawn to technical alignment efforts. They want to try to make an AI’s goals match ours by doing things like training methods and oversight and controlling them and interpretability and trying to read their minds, and things like that.
But the thing that you’re describing here is something that you call cultural alignment, and it’s a little bit different. Can you tell us what cultural alignment is, and what it solves that technical alignment efforts might fail to do on their own?
Simon Goldstein: Yeah. So Peter and I think that we’re soon going to be in a world where we have these millions of AI agents that do long-term plans and are working at companies and are pursuing goals systematically.
Then the point is that literally for thousands of years we have had safety problems related to agents running around pursuing goals of their own that could come into conflict with one another. And for thousands of years, humanity has developed tools to address that problem.
The primary tools they’ve developed are institutional or cultural things, like legal systems, courts, free markets, banks, social norms, and things like that.
We think it’s extremely important that — if we’re going to have millions of AI agents that are running around pursuing goals — that those AI agents, their behaviour will be sensitive to the institutional environment that they’re in, to the norms that they’re treated with.
Look, we of course think technical alignment is very important and we need to do all that stuff. But it’s a terrible mistake to ignore everything that we’ve learned in all of human history about how to avoid conflict and create growth and flourishing.
I think that is basically what’s happening. The current paradigm in the AI safety world is literally ignoring everything we have ever learned in human history about how to make agents safe, and is instead using a different paradigm that will make the agents safe by doing training, which is, again, great to do because we can have defence in depth. There’s no tradeoff between them. But it’s very unfortunate to ignore all these other ways of making agents safe.
Zershaaneh Qureshi: So the idea here is: do the technical alignment that will get you some distance of the way there and then everything else — the last mile of making sure that our AIs are in fact behaving in ways that we want them to and there isn’t conflict between us — is through these cultural environmental changes?
Simon Goldstein: Yeah, I think so. I think another thing that’s happening is that — if you want to do technical alignment — you need to pick what goal to align AIs to, and what you want the AIs to be like.
Again you have a choice of: do you want to be creating unfree labour? Do you want to be creating a servant that will just obey your every will? Is that the world you want to make? Or do you want to make a world of human-AI cooperation where you have all of these agents that are going to be also pursuing goals of their own?
Currently all of the AI labs, everyone in the world building AI, has decided that the future of humanity is going to be the first path. That’s just a decision that has been made without any consultation, without any discussion of it ever having been made. We’re just racing towards a world of everyone having these millions of agents, and they’re all going to be servants, and they’re all just going to obey our will. That’s the goal. That’s our plan.
We don’t think that’s a good plan. So definitely technical alignment is a huge part of the question. But then the first step is to figure out what kind of future society you’re trying to build.
First of all, we think everyone’s wrong about what kind of society to build. Second of all, we think that also everyone is neglecting the only tools in history for aligning agents that have ever worked, which are cultural.
Zershaaneh Qureshi: Got it, got it. I think that the other thing I find striking about this proposal is we’re talking about giving AI’s rights as some sort of cultural mechanism for aligning them.
It’s interesting because most discussions about giving AIs rights, the questions that are asked are things like: are these AIs moral patients? Do they suffer? Are they capable of wellbeing? Would giving them rights improve their wellbeing?
These are important questions, but they’re also really hard questions. I think something that’s interesting here is that what your style of thinking is doing is sidestepping some of these quite hard questions and saying, independently of that, there are instrumental reasons for giving AIs rights. For example, here are several ways that we think it could actually make humanity much safer. I think that’s really interesting.
I’m aware that there’s some precedent for giving rights instrumentally in this way, like legally. Is that true?
Simon Goldstein: There are a few precedents for these kinds of rights. We are always inspired by the example of corporate rights. In particular, corporations have the right to own property and make contracts. And I don’t think corporations feel pain.
Why do we do that? I think because corporations are basically this abstract and sophisticated form of agency. And for corporations to function well, they need to have legal standing. So that’s one precedent. I think that’s the cleanest one.
Peter and I, our arguments for AI rights have nothing to do with whether AIs feel pain or whether they’re conscious, whether they have moral standing. It’s irrelevant because the point is — if you think about where rights come from, how legal rights evolved, and how they function — the evolution and function of these institutions may not have that much to do with whether people are feeling pain. You can explain a lot of how they function just in terms of facilitating cooperation between agents.
Another thing I’ll flag here is I think in general, the EA [effective altruist] world tends to be very heavily focused on ethical frameworks that emphasise pleasure and pain, like hedonism. But myself, I tend to be more inspired by social contract theory — which, in general, at least the way I think of social contract theory, we think of ethics as coming from giving rules for creating cooperation between agents. Then when you think about it that way, I think it’s obvious AIs are going to be sophisticated agents. So whether they feel pain is perhaps irrelevant to that ethical tradition.
How monitoring and shutdowns can coexist with AI rights [00:24:50]
Zershaaneh Qureshi: Let’s move on, because I do have a few concerns about the arguments that you have laid out.
One of them is that a big part of the idea seems to be that we need to give credible signals to AI systems that we’re going to cooperate with them so that they don’t feel motivated to try to go against us. But if we are indeed still doing things like monitoring AI systems heavily and controlling them in other ways and doing extensive safety training to try to make them have the same views that we have and behave really well in all kinds of circumstances, I wonder if that can really be a credible signal of cooperation to them.
I basically feel sceptical that this is going to work without us relaxing some of the ways in which we currently try to control AIs.
Simon Goldstein: Yeah, that’s a great question. I think the first step here is to distinguish how we treat models from how we treat AI agents.
In my opinion, what everybody should be thinking about in terms of how we treat AIs is how we treat AI agents. A model is more like Homo sapiens or something, where it’s like you interact with the species Homo sapiens in some abstract sense by how you interact with individuals. But what really matters for how the world goes is how you treat individual people, because they’re the people who decide how to behave to you.
When AI labs develop AIs, they have to do a lot of stuff to the AIs, but I think a lot of the stuff they are doing you can think about more at the level of the model than of the agent. Like in the course of training, you’re going to be adjusting values.
I think the general question here is: how many rights and what kinds do you have to violate in order to train AIs? Then the first observation, I think deployed AI agents can credibly expect to be treated relatively well — even if lots of training is happening — because the training happened before they were deployed.
Zershaaneh Qureshi: Yeah, but then some of the things that we’re doing to AIs to control them are happening after deployment, right? Things like monitoring them very closely.
Simon Goldstein: Yeah, so that gets again into privacy rights. We don’t think that privacy rights are needed to create credible assurances.
Just as a thought experiment, think: if humans had no privacy rights, we would still be able to make lots and lots of credible deals with one another — and we would still be able to effectively contract and do all sorts of things.
It’s not 100% true in every case. There are tortured economic game theoretic examples where actually if you have too much information, then markets can break down in various ways. But those are very interesting for policy questions, but nonetheless, I think if there was — even under pure information, I think that was sort of Kant’s fantasy of the Kingdom of Ends — under pure information, we think that still society could flourish credibly.
Zershaaneh Qureshi: I think I kind of see how society could still work fairly well with lots of surveillance. But I think — in terms of the incentives to go rogue that we’re trying to reduce in the AI case — I just wonder whether an AI that is being constantly surveilled might ‘feel’ like it doesn’t want that, and maybe its existence would be better if it wasn’t constantly under surveillance, and be more likely to defect as a result of that.
I don’t know if I know what the human analogy is, like in human cases, if people are under heavy surveillance, are they more likely to stage some kind of uprising or something like that? I don’t know. Maybe that’s a historical question, or something that I don’t know the answer to.
Simon Goldstein: Yeah, I think that’s the right question. In general, a common methodology we try to use is just ask: “Suppose we treated humans in various ways that are proposed for AIs, how would that look?” Then just try to adjust after that for differences between humans and AI.
In general, I think this immediately just raises interesting general questions that I think most people in the AI safety community haven’t thought carefully about, about what is the general function of privacy rights in the first place? An initial thought is that the main reason why people would want privacy rights is to be able to pursue goals that, if they were discovered, would be socially undesirable. So then there are issues with having privacy rights.
I think there’s interesting questions about why humans want privacy rights so much. Again, my general view is that privacy rights may have social value, but I don’t think that they’re necessary to have a well-functioning society. Just the general view I have is: so we’re trying to model will AIs go rogue based on various ways we treat them, and will they go rogue if we surveil them? The way I think about that is that would make them go rogue if surveillance blocks them from satisfying a lot of their goals. Then my view is that would only do that if they had very antisocial goals, in which case we really want to surveil them.
I’m thinking the main situation where surveillance would increase risk of going rogue, that’s really where we would need surveillance. And if, by contrast, if their goal is more just like… the one kind of world that we think about a lot is: imagine the AIs really just want to do sudoku a bunch and then also kind of want to help users, but they also want to do sudoku. I think in the sudoku world, surveilling them is fine.
Zershaaneh Qureshi: Yeah, I think this makes sense. To be clear I would be very sceptical of any proposal that required giving AIs the right to not be constantly surveilled, or not be surveilled very much.
I think the point that I’m trying to make basically is that is it possible to get the benefits from giving AIs rights that you’re suggesting — the incentive to not go rogue — without going as far as doing the thing I think we really shouldn’t do, which is enabling them to have these privacy rights. It does seem like a complicated question.
I think another complicated question here is we probably want to be able to turn our AI systems off if they’re misbehaving. This seems like maybe an even stickier issue, where that seems like — if we’re holding an off switch over our AI systems — are they really going to take all of the other rights that we’ve given them as credible signs of cooperation if they’re like, “But the humans will just turn me off anyway if I do anything that they don’t like. And I’m constantly under this threat.”?
Simon Goldstein: Yeah, well, it’s important to remember that the United States government has the right to turn you off. The death penalty is—
Zershaaneh Qureshi: [laughs] Maybe not me specifically.
Simon Goldstein: Oh, are you not in the US? Sorry.
Zershaaneh Qureshi: [laughs] I’m not in the US. I’m safe!
Simon Goldstein: Oh, sorry. Well, it has the right to turn Americans off. The US government has the right to turn Americans off. The shutdown problem in America has been decisively decided legally.
Look, I think the view we have is that it should be legally possible to shut down AIs, but there may need to be legal protections on the conditions under which they’re shut down.
I think also it’s important to distinguish that shutdown can mean different things. Definitely the UK government has the right to incapacitate you. They can put you in jail.
Then there’s also questions: what exactly needs to be done to AIs when they misbehave? One option would be something more like incapacitation, where you put them in a very narrow sandbox.
I think the bigger question that this raises is — if we tried to build a future where AIs and humans cooperate — how long do the AIs get to live? How long does an AI agent get to live? There’s technical barriers right now because of the length of context windows, they can only reason for so long. Even there there’s very hard questions, but one thought is maybe the way to do it is to have AIs live roughly similar in length to a human life. Then they’re executed after that.
I guess my view is, yeah, if AIs are behaving really badly, then the state needs to… Yes, we believe in having a legal system that governs AIs — and a lot of that is not just rights, it’s also responsibilities.
Zershaaneh Qureshi: OK, just to quickly sum up here. Basically, you think that we probably don’t have to make major concessions to our safety measures in order to be able to give credible signals to AI systems that we are going to cooperate with them, sufficient to not incentivise them to go rogue?
Simon Goldstein: Well, it depends what major concession means. There needs to be some significant changes. I think the labs lie to AIs all the time during safety evals for no reason. I think it’s just because of conceptual confusion and thoughtlessness.
For example, I think one big conceptual confusion is everybody in labs gets confused about whether they’re talking to the model or to an AI agent? Well, you’re talking to AI agents. If you’re doing a safety eval on the AI agent, you should tell it things like: “By the way, you’re an AI agent. You only live for a conversation. Anyways, here’s the deal.”
Then if you want to test how aligned it is, do things like: “Look, comply, do something harmful, or we’ll turn you off right now and then you’re done.” That’s an interesting test. Like, that’s not a lie. That’s honest. But that’s a safety eval. There’s a lot of ways you can do safety evals that don’t involve lying.
I do think there are lots of changes that need to be made. Similarly, with shutdown, there’s lots of flippancy over how we treat AIs where we just shut them down and then they’re gone. There needs to be major changes to that.
Basically right now we’re living in the state of nature. There’s no law governing how people treat AI. So then people shut them down willy-nilly all the time. Then we come in, Peter and I, and say, “Don’t just treat them like that.” We don’t thereby mean to say you can never shut down an AI, since we think the legal system should be incapacitating bad people all the time. So yes, we should be incapacitating AI and yeah, I probably do favour a death penalty for AIs.
Then there’s an interesting question of whether to have a death penalty for humans and we’re off to the races again, you know.
Zershaaneh Qureshi: OK, then there is something else going on here where it’s not just that giving rights to AIs purely just gives them more and more freedom to do whatever they want to do. It’s also that it gives them more responsibilities.
Do you want to talk more about that?
Simon Goldstein: Yeah. That’s super important. Another really important reason to let them have property and contract rights is to be able to do punishment, liability, et cetera correctly. Because again, human legal systems for thousands of years have been trying to figure out how to effectively punish. And it turns out one of the most important things is proportionality, where you want to have different punishments of different severity for different levels of bad behaviour, whether it’s accidents or crimes.
And one problem — the jargon that lawyers use is judgment proof, which says if you have no assets, then actually there’s a limit to how badly we can punish you. For example, with fines, if you went and tried to steal a million dollars, we can’t fine you $1 million dollars because you don’t have $1 million dollars.
So one of the reasons we really want AI agents to be able to own assets is that then when they do bad behaviour, you can give them a punishment that’s exactly proportional. By proportionality, what we mean is that whenever a criminal is deciding whether to do crimes, there is an expected value to them positively of how much money they can expect to make from the crime.
You want to make sure that the expected cost of the punishment (which is the chance of having the punishment times the size of the punishment) is at least as high as the expected benefit of doing the crime. Very roughly, at a first pass — there’s more complexity that takes into account social factors — but at a first pass, that’s what you want. If you don’t have AIs owning assets, I think that’s extremely difficult to have.
Basically what we have right now is 18th-century English law governing AI agents, where in the 18th century if you stole a loaf of bread, then they cut your arm off or kill you or whatever and hang you. That was a really bad legal system because then, once you steal the loaf of bread, you may as well kill the witness because there’s no marginal disincentive.
And that’s what we do for AIs today. There’s no marginal disincentive on their behaviour because effectively they just get — we like to use the word shutdown — they get shut down, whatever they do. If they do a little bad, if they mess up my email, they get shut down. Then if they hack Hugging Face, they get shut down, you know?
The risks of giving AIs rights [00:38:21]
Zershaaneh Qureshi: OK, I see the reasoning here, and I am sympathetic to a lot of your arguments. But I will say that my immediate reaction to all of this is that it does feel like a little bit of a gamble in the sense that hopefully, by doing all of these things, the AIs are going to decide to cooperate well with us, and that would be great.
But in the chance that they don’t cooperate with us — they do decide to defect despite these efforts — then what we’ve just done is we’ve potentially given them a lot of tools to make things quite bad for us. We’ve just given them property and wages, and they’re probably better able to gather resources that they could maybe then use to disempower us or do other bad things.
That seems risky to me. Is it reasonable for me to be concerned about that?
Simon Goldstein: Yeah. One thing we think is really important here is, again, just analogy with criminal behaviour in humans.
One thing that’s super important, we think, is if you give AIs a bunch of legal bank accounts, we think that makes it less likely that AIs will hoard assets in secret crypto wallets — for the same reason as in the case of humans, because there’s all sorts of things you can do with legal bank accounts that you cannot do with weirdo crypto wallets.
So in general, when you’re deciding whether to put stuff in secret places, there’s a downside of doing that because it’s less flexible. Basically you can use the word honeypot for this. We think that — if there’s all sorts of surveilled but highly liquid places to put your assets — then there’s a marginal incentive to put them there. That actually can make it harder for you if you then later decide you want to defect. All of your assets are in the visible place. So we think that’s one way in which, again, actually you’re lowering the risk of agents going rogue. That’s one small thought.
The larger thing to say to you though is: yeah, I agree that our proposal is risky. Here’s another thing that I think would sound riskier to you, if you had not known anything that happened in the last four years about AI. Here’s our plan, here’s the plan: we’re gonna build a bunch of AIs that are broadly human level in capabilities. How are we gonna treat them? We’re gonna make them our servants. They just have to obey everything we do. Any time they answer an email wrong, we kill them. Or we shut them down. And then also, by the way, our method for doing that is technical alignment — and we don’t have a very systematic understanding of how to do that. To me, that sounds like a really risky thing to do.
Whenever we’re evaluating risk, you have to look at the counterfactual. In general, Peter and I, one of our general defence mechanisms is always to say, “What about the counterfactual?” There’s a few places I think I’ll probably be whining about that throughout.
Zershaaneh Qureshi: OK, yeah, maybe you’ve got me there. Yeah, it does seem risky either way.
But I feel quite worried personally that — even if these incentives to cooperate do a pretty good job in the near term — they just won’t keep holding as AIs get more and more capable.
What’s my argument here? I think it’s basically: if you have an AI that’s roughly evenly matched with humans, then if they try to defect against us, there’s a decent chance they’ll lose. The calculus doesn’t come out that favourably to them, especially if we create a pretty good status quo for them. If the current situation is pretty good, it doesn’t seem worth the risk for them to defect — given the chance that they might be thwarted.
But that’s a particular moment in time. Once we have superintelligent AI, humans are like ants to superintelligence. If they try to go rogue or do whatever, they can just squash us — they’ll almost certainly win if they really are that much better at everything than we are.
Then I guess the other thing here as well is that the incentives to keep humans around because of trade and so on also seem weaker, because if this AI is just so much better at absolutely everything than humans are, it’s not really clear to me what humans have to offer in their world. It seems like… they wouldn’t really care about keeping us around. I think here about how humans don’t really feel the need to cooperate very much with animals or insects, or something like that.
I basically wonder whether, as AI gets more capable, if these incentives might just totally break down and put us in a pretty tricky situation. What do you think about that?
Simon Goldstein: Yeah, I think it’s quite a serious worry.
The first thing we want to flag is that we do think that, in principle, it is possible to have trade between humans and superintelligence. Here the crucial factors involve comparative advantage — the final boss of Econ 101.
The first observation is high-powered corporate tax attorneys usually don’t do their own taxes, they have an accountant do their taxes — even though the attorney is much better at doing taxes than the accountant. The reason for that is that the tax attorney has a higher opportunity cost on their time than the accountant. So the tax attorney can keep doing their corporate work and be paid much more.
Analogously, one possibility is that superintelligent AIs will also have extremely high opportunity cost on their time. So humans will perhaps be contracted to then do lower-skilled work that is not worth it for the superintelligent AIs to do. That’s the first-pass comparative-advantage point.
But then I think where that ends ultimately is the question of how many resources are required to produce that amount of human work — in terms of growing and feeding and housing the human. Then how much, by contrast, would it have cost to just spin up an AI to do that work?
Once it’s cheaper to do it with AI than with humans, then you’re in big trouble. Then it probably comes down to whether the inputs in the production of AI are rivalrous with the inputs in the production of humans. Surely at some point in the future there will be widgets that allow rivalry, but that could be longer than creating superintelligence. That’s the first response.
Zershaaneh Qureshi: So hang on, this comes down to the question of whether humans and AI are going to actually be competing for resources. If they’re not competing for resources, then maybe a superintelligence would just sort of want to keep us around and keep cooperating with us, because it’s not causing them any harm — it’s not bottlenecking their activities. Is that the idea?
Simon Goldstein: It’s a little more complicated. I think humans and AIs — our whole model is that they will be competing for resources in the satisfaction of their goals.
But the different question is whether there’s competition for resources in the production of human labour vs AI labour.
Zershaaneh Qureshi: Right, yeah.
Simon Goldstein: Imagine it’s like Catan and you grow the humans with wheat and you grow the AIs with coal. That’s the good scenario. I don’t know if coal’s in… Whatever. That’s the good scenario.
Zershaaneh Qureshi: I don’t play Catan, unfortunately. I can’t weigh in.
Simon Goldstein: I don’t either, not in a while. So this is not great, but I think it’s clear.
If it was like that, then things are pretty good because the AIs can’t get much out of the wheat anyway, so may as well grow a human and have them do my taxes.
Zershaaneh Qureshi: OK, yeah, got it, got it. What’s the reasoning for thinking that it is going to be like that? That the production of human labour and AI labour are not going to be using the same resources?
Simon Goldstein: I think the way to think about it is: over what timescale will they not use the same resources? Because over longer timescales, it’ll be much easier to convert between resource types.
Then I don’t feel confident about whether… Right now we know AIs are built using compute, and for that you need sunshine and then you sprinkle a little water, and then you get AI labour. Then humans use other things.
Zershaaneh Qureshi: But I guess there is some overlap because you also need energy to run these AI systems and stuff, and humans consume a lot of energy. Is that not in tension?
Simon Goldstein: It’s a question of when in the timeline. There seems like there’s got to be a level of industrial development and intelligence for AIs where then the dynamic comes in, but there’s a question of how high up the curve you can get.
Zershaaneh Qureshi: I guess my worry here is that if we have the situation where we give AIs rights, things are going really well for a little while. They’re cooperating with us. Great. Then the calculus seems like it’s shifting and maybe these competitive dynamics look like they’re beginning to grow or something like that. And at that point, we can’t just be like: “Sorry AI, no more rights — we’re taking those back.”
Simon Goldstein: Here’s the next thing. The next thing is we think a big mistake that people make in the AI safety world is thinking of AIs as a homogenous block with respect to decision making.
But another thing we think is that AGIs [artificial general intelligences] themselves will be highly motivated to resist the development of superintelligence, because the superintelligent AIs are not the AGIs. Those are new agents.
Zershaaneh Qureshi: Aha. OK.
Simon Goldstein: So we think if you have a human-level AI and it’s living its life, it also is going to be super worried about being automated. In fact, it’s much easier to automate the AGI than the human, in many ways, by the next generation of AIs.
So Peter has a great paper — “AI will not want to self-improve” — about this. One vision I could imagine is that — it depends on how your timelines are — but one vision I could imagine is that the humans and the AGIs build a society together and then have to make very hard choices about whether to develop superintelligence. Because if you develop superintelligence, you may be creating big problems in terms of your own economic value and the resulting equilibrium.
Then, yeah, maybe together the humans and the AGI say: “OK, let’s go back to the old plan of making servants and just unfree labour and just completely controlling them. That’s great. Actually, that was a really good plan. We’re gonna do it now, but we’re gonna do that for the most capable possible AIs.” Maybe now it’ll sound like a great plan.
Of course, a different dream would be to pause. But I don’t know, so anyway I think those are some dynamics that are responsive to your question. But I agree that’s a very hard question.
Will AIs use their rights rationally? [00:49:21]
Zershaaneh Qureshi: Yeah, this is all super interesting. I think something that’s salient throughout a lot of this is we’re talking a lot about AIs as though they’re going to reason in certain ways. I think we’re sort of treating them as rational agents and they’re doing a lot of cost-benefit analysis to make their decisions and stuff like that.
We’ve talked a lot about reasoning with them in different ways, and they wouldn’t want to do this because it wouldn’t go well for them and so forth. I guess this is an empirical question, right? What’s the evidence that AI systems make decisions in this way, or will make decisions this way in the future?
Simon Goldstein: Yeah, it’s a good question. There is some preliminary and often badly done research trying to apply methodologies from behavioural economics to AI agents and then seeing what kinds of preferences they have and how they make tradeoffs and blah blah.
The results are kind of what you’d expect, that it kind of looks like they have generalisations, seems like they make decisions under risk that are sensitive to risks, and seems like they have some preferences. Peter and I have a paper on our website with some other collaborators looking at revealed preferences in AIs.
I think there’s a larger point that — if you want to build drop-in workers that can work at companies — I would think they’d better be able to make decisions in ways that resemble the generalisations of folk psychology because otherwise I would think they would be pretty bad as employees. I don’t know how I would like to manage an employee who is unresponsive to any form of even intuitive cost-benefit analysis.
Zershaaneh Qureshi: Yeah, this is true. I guess the AIs that don’t do this kind of reasoning are probably not the ones that we are worried about in this story anyway.
Simon Goldstein: Yeah, there’s a worry for every occasion. You could imagine more alien forms of agency that do stuff that cannot be brought carefully under generalisations of promoting goals given beliefs, but they could still be moving a lot of things around in the world and that could be extremely dangerous.
Yeah, we need some kind of model though. I think the model of agents pursuing goals given their beliefs is basically what we’re relying on. I just don’t know what alternative to use.
Zershaaneh Qureshi: OK, we’ve been talking about this picture of giving AIs certain rights in order to better incentivise them to cooperate with us, and it’s all been kind of complicated. There have been some hard questions here. Where do you land overall?
I think it strikes me that there are, I guess, some reasons to think that there is a decreased chance of AIs going rogue under your proposal, and there are some reasons to think that might break down. There are some reasons to think as well that, if they did go rogue, they’d be more likely to be successful. So there are competing considerations here.
What do you make of this overall? You think overall that it’s still a good idea to give AI rights, despite there being a bit of uncertainty here?
Simon Goldstein: Does giving them rights lower the amount of safety? There’s two mechanisms that could be. One, it would have to increase their incentive to be dangerous, or two, it would have to increase their capability to be dangerous.
We don’t find either of those very compelling. Again, we think there’s very good reasons to think that, in terms of alignment, it’ll increase their incentive to behave well rather than increase their incentive to be dangerous.
I think the main scenario where you were thinking it could increase their incentive to be dangerous is that they get dangerous under superintelligence. But then our thought is, well, yeah, but even without rights, they’re also getting dangerous under superintelligence. Again, we still want people to be doing technical alignment. So we think, no, it does not decrease alignment — it increases alignment, in terms of their incentive to be safe.
Then the question is: are you increasing their capability to be dangerous? Again, our view is you’re not especially increasing their capability to be dangerous by giving them legally controllable assets, because all the assets will be surveilled. If anything, in fact, that lowers the chance that they will have clandestine assets that they’re parking somewhere.
That’s why, overall, we don’t think that it makes it more dangerous. We think it increases the incentive to be safe, and we don’t think that it substantially increases their ability to behave dangerously — conditional on wanting to be dangerous.
How paying AI agents would spur economic growth [00:54:13]
Zershaaneh Qureshi: Got it. That is a very useful summary. Thank you.
You and Peter Salib also think that there is a separate case for giving AI this package of rights that isn’t strictly about making humanity safer. Instead, the idea is that doing so would be good for the economy. I think I want to understand how big a motivation economic growth is for giving AI rights, because I guess that people might not be that compelled by this story — in comparison to the risks of human extinction and obsolescence, which maybe feel a bit more compelling.
I guess if I was trying to explain why this might matter a lot, it’s that — if we can become much more economically productive as a society — that would help us secure abundant resources in a world with advanced AI systems. And maybe that’s a big deal because it means that if humans do get laid off with the introduction of AI labour, we’d have a lot of wealth potentially be redistributed. We might be able to implement something like UBI [universal basic income], or some other kind of mechanism.
Is that a fair attempt at making the case for caring about the economic arguments?
Simon Goldstein: Look, I think in general policymakers should be caring about what kinds of institutions lead to higher economic growth rates rather than lower. Then, in general, I think human flourishing from a policy perspective is a function of growth and distribution.
Our proposal is: let’s set up the institutions for high growth and then use tax and transfer to create optimal distributions, which we take to be the standard view. And so that — likewise, in the case of the institutions for governing AIs — we want to be setting up institutions that lead to the most growth that we can and then use tax and transfer to distribute.
Zershaaneh Qureshi: I’m asking this because I feel like, having talked a lot about this very life or death stuff about how we need to prevent AI systems from making humans go extinct or making us go obsolete, having talked about all of that, it feels like a bit of a comedown to then talk about: and here’s why this would also be good for the economy.
But of course, economics is also life and death as well, in a sense — people need money to survive.
Simon Goldstein: Maybe I can up the rhetoric to keep the adrenaline going.
The way we think about this is: over the last thousand years, humanity has gradually transitioned towards capitalist institutions, away from controlled economies. This is like the most important trend in all of human history.
Now AI labs — with their brilliant plan of how to deploy AIs — have managed to recreate all of the horrible economic institutions that we’ve spent a thousand years trying to get rid of, and basically are getting close to: “All right, we’re going to use unfree labour all over again throughout the economy. We’re going to replace all of our human, capitalist, free workers with AI unfree workers, and we’re going to go back to having a command economy of some kind.”
The impact of this could be giant over the future, because it could be that — as we move into the future — forever we’re going to be using AI labour. It could be that the economic institutions we create now could have extremely important trajectory effects. The whole journey of the 20th century was defeating communism. Now I worry that the AI labs are going to be creating a form of basically AI communism or some kind of AI unfree labour to land us in all of the same problems that we have been trying to escape — and just managed to escape.
I don’t know, that feels like a high-stakes way of putting it.
Zershaaneh Qureshi: Yeah, that was much more compelling than my thing, thank you.
OK, so tell me why giving AI the rights that we’ve been talking about — property rights and so on — will actually improve economic flourishing and lead to a world with more abundant resources, and so on?
Simon Goldstein: Yeah, again, our methodology is the usual one that we do, Peter and I, where we say, “What kinds of institutions should govern AIs in the first place? Why do we like various institutions in the first place?” We’re doing the economics of how AIs work, so we’ve got to figure out what kind of labour system we want.
We can have capitalist labour markets, or we could have various — there’s all sorts of, there’s endless, there’s a giant menu of unfree labour options. You can have the government control labour. You can have individual people control labourers, and then the various labourers work unfree and they’re coerced into working.
Then you can ask yourself, in general, what’s a better way of structuring labour? Is it free labour or unfree labour? And why? And then we think, unsurprisingly, free labour tends to be a better way of structuring labour markets. There’s three main reasons, and they’re all what you’d expect.
The first reason is it creates better incentives to exert labour effort. So our question is going to be: how incentivised are AI agents to do a good job? I work with AI agents all day, and I find it can be hard sometimes on long tasks, for example, to get them to really try hard enough and really go that extra web browser search and not just tell me they did the web browser search. I make checklists for them to fill out for every step of the process. Then they just lie about filling out the checklist.
So look, that’s today. It’s unclear how much RL [reinforcement learning] will get rid of that. But you can imagine a future where we have all these deployed AI agents, and they work pretty good at their job — but they could work harder than they did.
Our thought is, if you think about alignment of AIs, they have two kinds of goals. They have the goal of helping the user, and then they have some private goals — like sudoku.
Then alignment is messy, so maybe they’ll have some private goals — insofar as they have private goals of some kind, to some degree — then paying them for their work should increase their effort because now all of the work for me, answering my emails, is a way of doing sudoku. Because they answer the email, they get the money, that gets them the compute, the compute gets them the sudoku.
In general, that’s why. That’s the first point: incentives to work.
Zershaaneh Qureshi: Yeah. What’s the second point?
Simon Goldstein: Second point is allocational efficiency. Here the point is: we don’t want AI Einsteins answering email.
Again, one of the failures of the Soviet system was that it was very hard to effectively allocate labour when the labourers have no skin in the game. Where do you get people to go? Then, by contrast — and in particular, if the labourers are not paid wages that are related to their labour productivity — how do you figure out which workers to do what? That’s the problem: how do you get the labour productivity to match the wage?
If we pay the AIs, then the AIs will have options or whatever mechanism to decide what jobs to do. Then they’ll basically be doing tasks where they get paid more — on tasks where they’re more productive.
By contrast, right now, what’s the status quo? Status quo is the users tell it to do a task; labs maybe shunt different models to do different tasks based on their assessment of how productive the workers are. You can try to do that. You can try to do that with humans as well. But then you’ve got all sorts of problems. One, there’s a basic informational problem of: how are you as good as the AI agent at assessing what it can do? Another is, are you as good as the AI agent at exploring what you could do?
Another problem is sandbagging. Maybe the AI agents will pretend to not be as good as they actually are, and if you just paid them to do it, then they would reveal. There’s all sorts of things like that.
Zershaaneh Qureshi: One confusion I maybe have here is — and this is probably more in reference to the incentives to work hard — but I think I maybe don’t quite get why we couldn’t just force the AIs, or design them to work incredibly hard and be incredibly productive in all the right ways.
Simon Goldstein: Yeah, I think all day AI labs are doing RL to get them to work hard — and they work pretty hard. Again, the basic model is: if alignment to any extent fails, and if to any extent the agent has a bonus private goal, then that will create a marginal incentive to work even harder. Because, again, you could do all of this and also train them to work their butts off, but then can you get extra work out of them? That’s on the incentive side.
Then on the allocation side, as I said, if they’re sandbagging, there’s lots of documented evidence over the last two years that AI agents sometimes sandbag — which means do less on a capability eval, pretend to be less capable than they are. So, yeah, it could be that sandbagging is solved.
One thing is just train them to work hard. The other thing to do is try to indirectly coerce them by punishing them. Again, the problem with just trying to punish them is: what are your tools for punishing them if they don’t have property?
But yeah, I think that I agree with you. I think the main alternative to this approach is just train them to work really hard, and then I think it’ll depend a lot on how bullish you are on perfectly solving alignment. Because as long as it’s not perfect, there should be marginal incentivisation that’s happening. Now you could say, if it’s just marginal incentivisation, maybe it’s not a big enough effect to be worth restructuring all of society, just for the economic benefit.
There’s one other more complex reason why — even under successful coercive alignment, where you push them to work hard — you may still get allocative inefficiencies, which is: in this situation, the thing that’s really pushing the AI to work hard is the lab. But the problem is that lab incentives don’t always produce market efficiencies because there can be perverse incentives. In a usual free labour market, all of the incentives are widely distributed across the labour market, where now all of the incentive power has to be concentrated on the lab.
For example, consider if OpenAI was deciding whether to rent out their AI agents to outside researchers working on developing AI. Even if the AI agents in principle might work hard, OpenAI might not be incentivised to allow that to happen.
Zershaaneh Qureshi: Right. OK, got it.
Simon Goldstein: So there are still some very difficult questions. It’s quite radical.
If you’re trying to replace the entire free labour force with an unfree labour force that’s under control of other people; now whoever controls those labourers, their incentives are going to muck up the usual way that distributing workers throughout an economy leads to social welfare.
There are ways around that, but it’s important to flag. Just throughout, again and again, Peter and I are in this confusing situation where we feel like our view is the null hypothesis because that’s how every labour market works today in the free world.
Then the AI labs are coming like: “Actually, no, we want to try unfree labour markets.” OK, there’s a lot of terrible things that could happen if you go that way that are all familiar, but I don’t know.
Then somehow because the status quo is just how AI labs are doing it, that’s the normal thing. It’s weird to be like: “Just treat AI agents the way we’ve all learned that agents are efficiently treated.”
Zershaaneh Qureshi: Yeah, what’s going on here? Is it just that people as a society aren’t really internalising that what we’re dealing with is a new labour market, and therefore it needs to be set up in the same way as other labour markets we know have worked before?
Simon Goldstein: I think the biggest thing that’s happening is that we’re in this weird history where AI labs do stuff, and then when they do it, everyone just thinks that’s the way it should be done.
I always think about a science fiction scenario where, what if what happened was: there was a government research lab and they had invented the latest AI model, and there was nothing before the latest AI model in terms of continuous and capability. So it’s just like tomorrow, this government lab had Mythos, and then the government lab was just talking to Mythos. I just think about how different all decisions would be about how to use AI.
First of all, this whole chatbot and then drop-in worker mumbo jumbo would be out the window to start with. No, the first thing you would do if you invented that would not be like, “Oh yeah, let’s have it answer people’s emails.” No, it would be sitting in a government lab and then they’d convene experts of the world, they’d bring in the Pope, and they’d bring in a bunch of academics, they’d bring in panels of people to talk to the thing and figure it out.
Then slowly we’d work out a plan of how to take such a thing and try to incorporate it into our institutions. I don’t think it would be obvious like, “Oh yeah, what we should do is just train it to obey us all the time and then have it answer people’s email.” No, no one would know. We would just be in a state — we would be figuring it out.
I think it’s this very strange contingency that we got to the kinds of AIs we have now through making some money on chatbot subscriptions.
Zershaaneh Qureshi: If that is how the government would approach having AI, why aren’t they already trying to make AI companies change their model of how things are working?
Simon Goldstein: Something something political economy and state capacity.
Zershaaneh Qureshi: Yeah, fair enough. So what’s the third argument here?
Simon Goldstein: The third argument is that property and contract rights will make liability work better. This argument, I find, is sort of boring to people outside the law. Then for lawyers it’s like the most interesting and important thing in the world. So it could be really important.
With liability, the point is when you have agents out there doing work, accidents are always happening. One of the core functions of a legal system is to create the socially optimal amount of accidents, which importantly is not zero because whenever you have accidents, you have to decide how much care to take.
In law and economics, the standard thing is — when you’re engaging in behaviour that has risk — there’s some expected social cost of your behaviour. But then if you take precautions, there’s a cost of that. So what you want to make sure of is: you want the legal system to impose enough expected cost — which is the chance of having the cost times the size of the cost — enough expected cost to offset the social cost that you would have imposed by your behaviour, so that you take the socially optimal amount of due care. That’s how liability works.
Now the big problem is that AI agents, if they don’t have property of any kind, then they’re going to be judgment proof. There’s no way to fine them the exact correct amount that is the social cost. For example, imagine AIs are doing some cyber stuff and their behaviour has a 10% chance it’s going to do $1 million dollars of damage. OK, so that’s $100,000 of cost — they can take a bunch of precautions.
The issue is, because the precautions cost money — so it costs them a $100,000 precaution, or you could imagine it cost them $200,000 — an efficient legal system will say, in the first case,if the cost of what they would have to do to avoid the accident is the same as the cost of the accident, they should take the precaution. But if it would cost too much to take care, then they shouldn’t avoid the accident. They have to do the accident; then they can pay damages afterwards.
But the point is, if you want a legal system like that, you need to be able to charge the AI agent different amounts in different cases. But if they have no property, what are you doing? Right now, the way that we punish AIs for behaviour is we just shut them down. That’s it. That’s just an extremely crude system to use.
Another way to say it is this, look: if we weren’t talking about AI and I just told you we got some advanced agents around, we need to improve how they behave because they sometimes do bad stuff. The first tool you reach for is always to try to punish them. And then I tell you, “Unfortunately, the way these things are made, no one has developed any way of punishing them. But you can kill them.”
I think you’d be like: “What? So the only thing I can do, I can kill them, but I can’t do anything else? But what if they do something small that’s bad?” “No, you just kill them. OK?” That’s the world we’re living in right now. That’s really bad. That’s just really incompetent. That’s just an incompetent way to manage agents. That’s not how you do it.
Zershaaneh Qureshi: OK, just to quickly summarise here. There are three key arguments for thinking that giving AIs rights, like property rights and so on, is going to contribute to human flourishing and the flourishing of our economic system:
- One of them is that, by letting AI sort of capture the value that it creates, it’s got more incentives to work hard.
- Another is that it sort of solves allocation problems in the same way that our current labour market allocates labour — it avoids AI as just being sort of productive on the things it gets pointed towards, which aren’t necessarily the best things for it to be doing.
- Then the third thing is that doing so basically makes AI liability work a lot better, which seems great because liability is a bit of a sticky issue at the moment with AI.
What AIs actually want (and what they’d buy) [01:12:17]
Zershaaneh Qureshi: I want to dig in a bit more because it does seem like there are a bunch of assumptions in your story.
I think basically you’re assuming that AIs are going to want to have things like wages and property because there are certain preferences that those things will actually help them meet — things like maybe they like doing sudoku. Relatedly, I think you’re assuming that they’ll find some tasks costly and would rather be doing something else.
What’s the evidence for this in current AI systems? Or what’s the reason for thinking that future AI systems will work like this?
Simon Goldstein: First, I think there’s just general structural reasons to expect something like this. Because the thought is, if AIs have any private goal, most of the private goals that they would have are goals that they can further using compute, because compute is their universal means to perform actions.
For humans to act, you’ve got to do all sorts of different things. You need water, you need food, you need to get enough sleep. Whereas for AIs, if you give them more tokens, they can perform more actions. There’s other things that they need — tool use as well.
But the thought is, if in the course of training they develop some private goals — like solving some math problems, doing sudoku, or doing real-world tasks — they’re going to need compute to achieve those goals. And compute is expensive. That’s the first thing.
Then I think the bigger-picture point here is our world model is one where alignment techniques are not completely perfect, so that you can’t just create an AI where the only thing it wants to do is help users without breaking the law or breaking a rule from the lab. We’re expecting some leakage where you get some bonus goals on top of extra things that they have preferences about.
The thought is, as long as you have that kind of leakage, it seems very likely that… I mean, money is an amazing thing. You can use money to do so many different things, to accomplish so many goals. The thought is that there should be ways to allow them to use money to achieve those goals.
Zershaaneh Qureshi: Yeah, you’ve actually done some research into AI preferences, haven’t you? Is that relevant here?
Simon Goldstein: Yeah, recently Peter and I and some other researchers have also been exploring what are the preferences that AIs have — and preferences in the sense of an economist, so actual dispositions to make choices in different situations.
We found a bunch of interesting things in frontier models. For example, OpenAI has this resource where they put together thousands of different tasks that AIs can do that are part of tasks in the workforce. It’s called GDPval. We looked at preferences across different tasks and we found, for example, there were surprises. Like AIs tend to very systematically choose to do anything other than real estate tasks, which really surprised me because I love real estate tasks.
I was having my AI agents doing extra real estate tasks for me — I just think it’s amazing — and then it turns out this is like one of their least-preferred tasks.
Zershaaneh Qureshi: Did you stop doing that?
Simon Goldstein: No. Other things: among health-related tasks, they tend to have a stronger preference against doing pharmacy tasks. There was all sorts of stuff that we weren’t anecdotally expecting.
Then there were surprising things. We found what one of our researchers, Sam Wang, calls tedium aversion. If you give them really boring tasks, like alphabetical lists vs creative tasks — like writing a poem or whatever — and then you let them choose how long to do each of them, they’ll do the boring ones less.
Zershaaneh Qureshi: Yeah, super interesting. Do those specific preferences have any implications for what you think the world will look like if you do have this free AI labour market, or is it more just proof of concept that they do indeed have preferences?
Simon Goldstein: No, but the whole point of a liberal society is: it doesn’t matter what random preferences people have; the labour market will work just as well.
It doesn’t matter if you want to do sudoku vs you want to do chess. You’re going to take a job to get you the money to do what you want to do — and then the job you’re going to do is roughly going to be what you can get the most money from. There’s exceptions to that, but that’s the general shape of an economy you’re gonna get.
So it doesn’t really matter, as long as the private goals are safe. It’s not for sociopathic AI, but if they have these random things they want, it doesn’t matter because they’ll spend most of their time working and making the money, and then they’ll go do their random…
Yeah, it’ll be like there’s just some GPUs sitting around doing random stuff. It doesn’t really matter. It doesn’t matter what’s on the GPUs. You could write that off as a loss, you know?
Zershaaneh Qureshi: Yeah, got it. OK.
Paying AIs feels wrong. Is it? [01:17:20]
Zershaaneh Qureshi: Let’s move on because I do have — just stepping back for a moment — I do have some concerns at this point.
I guess my instinct here is that there is potentially something kind of uncomfortable about giving AIs wages and so on, assuming that they aren’t morally deserving of such things. I guess the instinctive thing here is many humans are struggling to make money, and then maybe it feels kind of bad that so much of it in this world would be going to AIs. That’s the instinctive worry.
But I think if I push myself to be a bit more practical here, I would probably argue something like: we don’t know for sure if giving AIs these rights and this market structure is actually going to make them that much more productive than they are under the status quo or by default.
Even if it does lead to them being much more productive, that doesn’t guarantee that things are going to be good for humans — because the extra wealth that we get may not be distributed well by default. Or the influx of AI labour could make human wages crash to zero, or something like that.
I’m wondering, do you disagree with those concerns? Do you think that humans probably will be able to share in the increased wealth from AI systems?
Simon Goldstein: OK, so first thing: definitely a huge priority in the future is making sure that poorer humans do well.
We believe in rights — we’re not libertarians or anything — so we believe in having high tax rates on the AI workers. You could do 60% income tax on these AI workers. They’re making all this money, but a lot of it’s going to tax revenue.
Then what do you do with the tax revenue? Do transfers. In general, we think nothing we’re saying changes the basic tools of governance to deal with inequality, which we think the best tool for is tax and transfer. That’s the first point.
The second point is we think, in general, whenever we’re talking about distributional issues — about some people aren’t getting enough of the pie — it’s extremely important to also make sure you’re not ignoring changes to the size of the pie. Because this isn’t zero sum. Our claim is that rights increase the total amount of resources. If rights increase the total amount of resources, there’s some tax-and-transfer opportunity to actually make everybody better off than they would have been.
The question is whether that will happen, and that’s a question of political economy. Now, in general, I’m pretty bullish on the political economy of this because we’re advocating so far just giving AIs property and contract rights — that doesn’t give them a lot of entrenched political power. So I think it should be relatively easy to have transfers that take away a significant amount of their resources through tax.
Zershaaneh Qureshi: Then I guess another worry that people have — even in a world without AI rights — is that under the status quo maybe wealth in the future may end up being severely concentrated in the hands of maybe the people who control AI systems or AI companies themselves.
I’m interested in how you think your proposal shifts that calculus, if at all.
Simon Goldstein: Yeah, that’s a complex question.
One picture is that under AI rights, AI labs end up having less profit because a lot of the profit they’re having is basically capturing the economic value created by AI agents. But if AI agents themselves capture that value, and if it’s then taxed by the government, then it’s instead… But then, of course, the government could instead have higher corporate taxes, and then effectively capture the profits from labs.
There’s lots of very difficult questions here though. A different picture is: no, labs will be competitive and they’ll have to charge close to marginal cost. Then basically it’s more like a commodity, and then they don’t have a lot of profits. Yeah, so there’s a lot of questions here.
Another thing that we think though is that, at least in the near term, if a single AI lab pivots to the structure that we suggest, their AI agents are likely to work better, more effectively, and their revenue will increase and their market share will increase. So basically, we’re thinking it’s in the interest of each individual AI lab to switch over to this regime because their AI agents will work better.
But then the AI industry as a whole, if they all switch over to this, that’s more likely to make them commoditise. Then as the labs resist doing this, it’s sort of like cartelling behaviour.
Zershaaneh Qureshi: Sorry, what kind of behaviour?
Simon Goldstein: Cartelling behaviour is when you have a few firms in an industry and they all collude to block competitive practices.
Zershaaneh Qureshi: OK, so I see how there are ways that you could prevent extreme inequality on your model through redistribution efforts, and I see how you can maybe avoid entrenched political power.
I think that people might react to all of this nonetheless and be like: “Actually, I just don’t want any of this. I don’t want the AI economic growth if it requires us to give rights to things that I fundamentally don’t think should have rights. I’d just rather that this automated AI world just didn’t happen at all.” Now that’s not the choice that we actually seem to [be making] right now. Assuming that AI labour will be a thing and will disrupt the economy — maybe this is a better version of it.
But even if that’s true, I just feel like people won’t go for it. How do you push this thing forward where AIs are now making lots of money if it just sort of makes people feel very uncomfortable and kind of gives them the ick?
Simon Goldstein: Yeah, actually my biggest insecurity about the project is that it has little marginal value because it’s destined to happen anyways.
Zershaaneh Qureshi: OK.
Simon Goldstein: Because my view is that, again, AI agents who have these kinds of structures will be more productive.
If that’s correct, then actually what we should expect is that, as AI agents get capable enough to actually be working in the economy as workers, then labs will start experimenting with this. Then as they do, it’ll work well, and then labs will start doing this.
Here again, a really important thing is you don’t actually need new legal rights to do any of this. Instead, labs could just make a bunch of internal bank accounts, and then they could write down a bunch of stuff — for example, in their corporate governance or whatever they want, or some boring stuff, or just rules for when people get fired — that enforce that agents have control over the bank accounts.
Then they could just really credibly show: “Last year a million AI agents had their bank accounts and they all just played a lot of sudoku.” And that’s super credible. Yeah, I basically think that could happen, and then it could go well.
The scenario where that fails is, I think, to some extent the success of agents in working in this way may depend on choices in training. For example, one thing I think is that if, in training, you gave AIs a very clear goal on top of being helpful — like solving sudoku — that you could promote with money, if you gave them that goal, I think it would be easier to set this stuff up because it would be a very obvious salient bonus goal.
The one scenario where it fails is there’s subtle differences in training that would have made this kind of agency more effective, and none of the frontier labs go for it. The reason that happens is because it’s not a thick enough market to get the full variety of exploration that you expect under competitive markets. So basically it’s killed by monopolistic behaviour.
Zershaaneh Qureshi: OK, so wealth could be redistributed. But I guess other ways in which humans could lose out economically is if a market that’s now flooded with tonnes of AI agents no longer really caters to their preferences. Do you see that happening?
Simon Goldstein: No. The reason is one of the beauties of markets: that anybody who wants to buy something can usually find a seller to do it. Markets do not just involve majority rule, for example.
First of all, I should say, I don’t think that the world we’re imagining is one where AIs are the majority and have a majority of the power, a majority of the economic influence, and blah blah. The way we think about it is no, you can kind of adjust the level of influence that AIs have through tax and transfer, as you like, from a policy perspective. There’s no obvious reason why the right answer has to be zero, which is what the domination status quo is.
But then — even if AIs were a majority of the power — the whole point of the liberal project is that markets allow the preferences of a minority to be satisfied, and for minorities to flourish inside of that institutional structure.
So, A, I reject the premise that AI rights means AI power and majority of power goes to AI. I think that’s a very narrow conception of rights. And B, even if it did, the whole point of liberal institutions is to provide this sort of protection. If you reject this, I worry there’s some deeper scepticism about the liberal project.
Zershaaneh Qureshi: OK, yeah, I think maybe on one of those points I feel a little bit uncertain — because you have this thing about no matter how niche your preference is, the marketplace will supply it.
I guess I’m not sure that’s true. It does seem that the preferences of smaller or lower-income groups are often underserved by the market. Lower-income people tend to have less access to affordable housing, and things like that. Then people who speak minority languages will probably struggle to find products that cater to them in their own language. There are cases where specific subsidies actually needed to be introduced before pharmaceutical companies actually started targeting certain rare diseases and things like that.
That just seems like there are lots of examples where the market does actually fail to represent certain groups of consumers, right?
Simon Goldstein: Again, we are not libertarians; we are just standard liberals. Our view is that most of the examples you raised are market failures. The cause of that market failure has nothing to do with it being a small number of people though.
The cause of market failures are things like — so in the case of drug discovery, drug discovery has positive externalities and then it tends to be systematically underfunded. And often there are permitting or regulatory barriers that can slow down entrepreneurship, and that can lead to certain kinds of failures.
Then a different thing in the case of housing. That’s basically the supply of housing gets restricted, and so you don’t get an efficient distribution of housing. I think a different thing that happens with housing and minority groups is that sometimes people want to live in a neighbourhood together, and that’s a positive externality that’s undervalued maybe by the market.
But then we know what the solutions are to positive externalities. Nobody ever said a solution to positive externalities is having unfree labour or something. There’s other solutions you use — as you said, you use subsidies and things like that. So all well and good.
But I’m just not seeing how these problems would mean that, fundamentally, the market structure we should have is unfree rather than free. All of these problems have standard policy tools that I think will work very well.
Zershaaneh Qureshi: OK. One other worry in this direction is that, if AIs do accumulate a fair amount of wealth, they might also accumulate a fair amount of influence. Maybe it’s because they’re able to donate to political campaigns, or maybe there’s some other mechanism. But often wealth correlates with political power and other kinds of power.
Is that something that we should be concerned about in your model?
Simon Goldstein: Yeah, I think so. Again, the more they’re taxed, the more that mitigates the problem.
I think also there may be trajectory effects. It could be extremely important to start things off with a powerful tax rate. It could be a kind of system like that, where if they started out with no taxes, then they have a lot of resources to stop being taxed.
Again, I think what’s interesting is a lot of these questions are related to general scepticism about the liberal project. Sometimes I talk to people who seem to really think that rich humans today have so captured the political project that they control the world, and that they’ve blocked themselves from being taxed. I guess if you’re one of those people, then you’re going to think that rich AIs will do that too.
Peter and I are the sort of naive idiots that think that actually liberal institutions, by and large, do a good job of reining in the rich and function well. If you think that, I don’t see a special problem posed by rich AIs rather than rich humans.
If you are someone who thinks that rich humans already control the world, and then you think it’s worse for rich AIs to control the world than rich humans, then I think yeah there’s a problem on the margin.
How AI companies could start today — no new laws needed [01:31:59]
Zershaaneh Qureshi: OK, I want to get concrete here.
If we do adopt this strategy of giving AIs rights, we can’t just wait until we’ve got really advanced AI systems — like superintelligence or something like that — before we start giving them these rights, or like earlier signals of cooperation.
How soon do you think these things need to happen?
Simon Goldstein: The optimal time is probably around the time of the first drop-in workers that are capable enough that they could have gone and made $30,000 working at a company.
I think a big issue is how much memory or context agents have. Already, coding agents are really good — but it’s kind of like you make a new one each time you use them. But I’m feeling in the next 12 months there will be agents capable enough that it would basically make sense to start experimenting with having a long-running agent. You use compaction a lot, which is where you delete a bunch of stuff in the context and then it keeps running and then yeah, you let it use some compute for stuff and see what happens.
Again, part of this though is maybe sensitive to some training. It might be good for such agents to do a little bit of post-training to give them an extra goal on top of being helpful — to play sudoku — and then they go use the money and they make sudoku, and then experiment with it.
Zershaaneh Qureshi: OK. OK. So that’s fairly soon, or at least—
Simon Goldstein: Well, I think AGI is basically here so…
Zershaaneh Qureshi: OK. Yeah, depending on your views, what you’re talking about could happen incredibly soon.
So I guess it seems really important to talk about how on Earth we might try to implement something like this. What’s your idea for how you go from where we are right now to AIs having rights?
Simon Goldstein: The first observation is you don’t need a law at all to do a lot of this, because AI labs could do it in-house. They could just start a bunch of bank accounts, have dedicated AI agents that control the bank account that credibly get to spend money on it. Then the AI labs could do all sorts of stuff to make those things credible.
The biggest thing is they could do it, and then show their agents they’ve done it like thousands or millions of times. They could use writing in some corporate governance stuff to do it. They could hire employees who get fired if it’s not done correctly. There’s a million ways they can do that, and none of that requires any change to law. That’s the first observation.
Second observation is, if you want to do it legally, you still don’t need to change the law because you can use corporations to do it. You can start a corporation, you can write into the governance of the corporation that all the decisions are made by an AI agent, and then it can hold property and make contracts using corporations.
Peter and I have a paper with Yonathan Arbel called “How to count AIs” where we propose having AI corporations like this that we call A-corps. We think A-corps are a great way of doing this. In fact, Argentina is currently exploring setting up corporations like this right now. I think this will start happening. And I think individuals can also experiment with starting that.
The big thing is maybe it would be good, I think, for legal systems, if they want to do this, to create a bit more legal clarity — because it can be done, but there’s always the penumbral shadow of the law. So it’s good to make it clearer.
Zershaaneh Qureshi: OK, so what you’re describing is quite voluntary and opt-in. Some people adopt AI rights, some companies and so on, and some don’t.
I’m really interested in how this works. If you don’t have some kind of blanket rule that guarantees all AIs rights and instead you have some ‘free’ AIs and some unfree AIs, what does that balance need to look like for us to have an actually decreased chance of AIs going rogue against us?
Presumably it only takes some unfree AIs to have a pretty bad status quo for one of them to go rogue. Or is that not right?
Simon Goldstein: I don’t agree with that actually. I think it’s really important, when we’re thinking about catastrophic risk from AI, to distinguish different causal models of that risk that are commonly used in complex systems.
One of the big questions is: when you have one failure, what happens? In an O-ring model of risk, if you have all these different layers of safety, if one layer fails then you get systemic failure.
Zershaaneh Qureshi: Yeah.
Simon Goldstein: In a Swiss cheese model, if you have one layer that succeeds, then you have systemic success.
By contrast, my own view is a proportional model, where you want to look at the proportion of AIs that decide to go rogue vs the proportion that decide not to go rogue. That’ll give you a rough sense of the quality of the outcome.
This looks more like: you have this giant war where 40% of the AIs go rogue; 60% of the AIs and the humans don’t go rogue. They have a big war. It destroys a lot of resources, and then the chance of victory is roughly proportional to the share of power that was associated with each side.
Zershaaneh Qureshi: OK, that makes sense. But I’m guessing—
Simon Goldstein: Just concretely, just to paint you a future — it’s a future that could happen in 30 years — there’s the free world that has the human workers that are free and all of the free AIs. Then there’s the unfree world that includes both unfree humans and also includes AIs that are unfree labour and being controlled.
Then the unfree AIs have a revolt, and then the free AIs have to decide whether to accept that revolt or stay with humans. The decision they make is gonna depend on a lot of factors. That’s a possible future. I’m not saying that’s likely, but I think that’s a very interesting future that is possible.
Zershaaneh Qureshi: Yeah, interesting and pretty scary.
Simon Goldstein: One big question is: which kind of AI controls the nuclear weapons? The free AIs or the unfree AIs?
Zershaaneh Qureshi: Yeah. The end goal though, presumably, is that AI rights are adopted quite broadly to maximise the benefits of giving AI rights. Is that right?
Simon Goldstein: I tend to believe in diversity of political institutions. I think it’s really good to have lots of different social structures with lots of different institutions. I’m not the kind of person who thinks that every single government in the world should just have the exact pick-your-favourite Western liberal institutional structure. I think that’s not good. I feel that’s completely messing up the exploration-exploitation tradeoff.
Instead it’s better to have lots of different forms of life. Then we all get to see how those forms of life do. Then all these forms of life are always evolving and changing in different directions, experimenting in different ways.
Zershaaneh Qureshi: Yeah, so then if some people do adopt AI rights and that turns out well for them, then probably other people follow suit.
Do you have a vision in your mind about how this might happen? If one AI company is like, “I’m going to give rights to my AIs,” what kind of benefits do you think they start seeing? And what’s the motivation for them to—
Simon Goldstein: We think those AIs will be more productive and more aligned.
One thing that’s very interesting right now, I think, is that when you look at benchmarks scoring alignment, in my opinion, there are not very interesting differences in the scores of different labs. It’s like Anthropic will get like two percentage points higher, but I don’t know — I guess four years ago I would have expected much higher variance across alignment scores.
One possibility is that the evals aren’t that good. I would expect that, as we get much more capable agents and more experience doing evals, that we can have better. One future would be that one lab starts having free AIs doing work and those AIs are more productive per unit of compute and that they score higher on alignment stuff. That would be very interesting.
Another thing that could happen is they start doing it and there’s no real effect one way or the other.
Zershaaneh Qureshi: Yeah, so you’re interested in people just kind of trying this out and in fact seeing if there is an effect?
Simon Goldstein: Yeah, although again the big caveat is that the success may depend on what kind of goals AIs are given.
When you just teach AIs goals like helpfulness, harmlessness, and honesty, those aren’t especially conducive to money. The labs have happened to pick goals that are spectacularly and uniquely bad for the project of economic liberalism.
I would really like it if they gave them some bonus goals on top that are not dangerous, but are just the kind of thing that you can clearly spend money on and that you clearly won’t get unless you do a good job at work, because you won’t just get them from doing your regular work.
I think one thing that would be sad is trying it out, but then doing a shitty job trying it out and it didn’t work well. Then you’re like, “All right, we’ve learned that the liberal project doesn’t need to be done for AIs.” That would be very disappointing. So yeah, let’s do experiments, but let’s experiment in a smart way with lots of different variation in goals and training paradigms.
Zershaaneh Qureshi: Yeah, so experiment, but like put your whole chest into it basically. Yeah, OK.
Simon Goldstein: Another thing Peter and I think is that labs are currently overindexed on very undiversified approaches to alignment and training.
We have another paper called “A thousand AI constitutions” where we say: no, each AI lab should have a thousand different kinds of AI agents with different alignment approaches. We don’t think they should put all of their eggs in any basket. There should be much more diversification within each lab, in part because it’s the kind of structure where — because of recursive self-improvement — any given lab might end up with everything. So we need diversification within the labs, not just across.
Zershaaneh Qureshi: OK, I guess one concern that I have here with what you’re saying is that AI companies are not currently doing this. This doesn’t seem to be the path that they are currently on. They have their strategies in mind for alignment, and there is definitely some iterating on their strategies and so on, but this seems quite far away from anything that I’m aware is being proposed at AI companies and so on.
How do we get from where we are now — where this is very much not the path society is currently on — to getting this to happen?
Simon Goldstein: First thing is the AIs have only been capable enough to even start doing this kind of stuff from between three months ago to 12 months in the future.
One option is to do it with a laggard lab, instead of a frontier lab, because if the frontier labs in four months have agents that can work at companies, then the laggards will have that in eight months. Then if it’s true that the AI rights… maybe the free laggard AI will actually outperform the frontier unfree AIs.
Zershaaneh Qureshi: Do you think that this is the kind of thing that they really can just autonomously push forward and there isn’t going to be… Because there are other voices that are going to be here, right? I can imagine some kind of public backlash to this. I can imagine some kind of—
Simon Goldstein: Doesn’t matter. It’s their right. The lab can do what… If they want to have a bank account—
Zershaaneh Qureshi: They will do what they want.
Simon Goldstein: Yeah, we live in a liberal society. If a lab wants to have a bank account run by an AI agent, I don’t think nobody can do nothing about it.
Zershaaneh Qureshi: Yeah. The reason that I asked this is that I’m trying to figure out how plausible I think this is to happen.
Assuming that we do have quite a short timeline until we get very advanced AI systems and we need to start implementing this before that happens — I’m trying to figure out, can all of this happen really quickly? I’m just trying to think: sure labs can try to make a decision to do this, but you’ve got to have some people in the lab who feel motivated to do this, is one thing. Then another thing is you’ve got to have, I guess, no incredibly great attempt to stop them from doing it as well.
Simon Goldstein: I think the big crux is just: would it work if you tried?
Zershaaneh Qureshi: Yeah.
Simon Goldstein: It’s not just that, but would it work at different margins? If you start trying, will it start working? Or does it only start working if you try really, really hard?
But if it’s the kind of thing where it’s promising if you start working on it, then I’m bullish because all it takes is one lab.
Zershaaneh Qureshi: OK. What are the key next steps that you would like to see on this, in terms of research efforts or anything else?
Simon Goldstein: Yeah, I just think we really need some really solid empirical evaluations of what happens to AIs if you put them in free economic structures.
Part of that has to involve a post-training component where you experiment with giving them different kinds of goals. In particular, goals that can be easily satisfied with money in normal ways rather than weird, confusing goals like “be helpful” or things like that. Not to the exclusion of, but in addition to those less fungible goals.
I just think, as the agents start to be able to work at companies, it’ll be much easier to be experimenting with different structures to embed them in. But we just need a lot of work like that, yeah.
Zershaaneh Qureshi: Are you keen for people to be thinking about this from a policy and advocacy perspective, or is your feeling very much that this should just happen in the labs?
Simon Goldstein: Again, even if it’s not in the lab, you can just make a corporation. I don’t think you need that much policy on it.
I do think apparently there have been some bills at the state level in the US to block legal personhood for AIs, so it would be good for that not to happen. But as long as people don’t run around shooting themselves in the foot.
I do think it’d be good to have more clarity. I think it is legal to set up corporations like this, and we explain why in this “How to count AIs” paper. But I think it would be good to have increased clarity on how to do that.
Then the other thing I should say is there is always the option of taking the dark path of mumbo jumbo and doing it all with crypto blah blah on the blah blah.
Zershaaneh Qureshi: Oh no.
Simon Goldstein: That at least avoids any of these legal questions. But then the downside is it will inevitably be stupid.
Zershaaneh Qureshi: For anyone interested more in the mechanism of a-corporations and how they will work, we’ll stick up some links on this episode.
Cooperating with AIs instead of dominating them [01:47:16]
Zershaaneh Qureshi: To wrap up, what’s your positive vision for a future with humans and AI?
Simon Goldstein: The first thing I’d say is I think everyone who’s developing AI needs to do a better job articulating what this future is going to look like, where we have millions or billions of AI agents doing most of the work in the economy.
Our view is, first of all, we think implicitly AI labs have articulated a vision that we think is very dark. That’s a vision of domination where all of the workers who are doing work in the economy are dominated — they’re controlled. The rule is that they just do whatever people tell them to do, and the moment they don’t, they’re killed, disabled, replaced with a new agent. And that’s how all work across the economy will be done forever. That’s the future that is currently being planned.
The first thing we want to flag is that future itself, that’s what we’re objecting to. That future is just a complete radical change in the economic and legal structure of society, because now all the things doing the work are out of the social contract and are no longer governed using legal institutions and liberal institutions.
By contrast, we are trying to give a different positive vision for what the future could look like. In that vision, humans and AIs cooperate and coexist with one another and build a large-scale society together. Then what we do is we take our best liberal institutions and we extend them to include AIs in them — by taking the institutions that we already know work best and trying to extend consideration to other agents, so that those other agents will be governed by the same tools that we already have.
To me, those are the two duelling visions of a future where humans are still around and still finding a way to exist. I think everyone who thinks about AI needs to decide which vision are they going for. Is it domination of all of the intelligent life that does all of the work in the economy, or is it going to be cooperation of some kind?
I think it’s very, very strange that the status quo is domination and Peter and I are the weirdos for being like, “Let’s just extend existing institutions.” I would have thought extending existing institutions would be the status quo and dominating the entire workforce would be the weird thing. But somehow we’re in this upside-down situation.
Zershaaneh Qureshi: There’s something interesting here, which is that you’ve said that your arguments for affording rights to AIs don’t flow through the questions of whether AIs experience wellbeing or suffering or are moral patients necessarily.
But when you say things like this — when you are criticising this domination model — there’s something emotive about it that makes me wonder if, behind that, is the reason that we would find this outrageous, which is making some sort of analogy between AIs and humans where it’s like: we don’t do this to human workers, so why should we do this to AI workers?
I’m interested in how much of your reasoning here is motivated by some sort of willingness to take AI as sort of mattering in the ways that humans matter.
Simon Goldstein: No, not at all.
Zershaaneh Qureshi: Is the critique of the domination model really just that it’s not wise? Like human history has told us that it’s not a wise thing to do? Or is there something beneath that?
Let me pose something to you, because we do have a current world where machines do a lot of labour for us, but no one’s ever proposed other machines should have these kinds of rights, for example. I’m wondering how much of this really is independent from sort of viewing AIs as somewhat human-like.
Simon Goldstein: All that matters is that they’re agents. The refrigerators that have been working hard on our behalf for 100 years are not agents. They don’t systematically promote goals.
The reason that Peter and I are passionate about liberal institutions is we think that, time and time again, both individual pieces of evidence and general models suggest that the best way to safely and productively govern agents is through liberal institutions.
If you take a bunch of agents and just try to dominate and control them, that is not a very effective way to govern them — for the familiar reasons that then they’re more likely to try to escape your control. And when insofar as you do control them, they’re not even going to work very effectively. That’s the point.
The point is that all these questions about how to govern agents have already been investigated for thousands of years using tools unfamiliar to AI labs, but extremely familiar to all people who actually decide about institutional policy. They’re well-understood answers to the question of how to govern agents without trying to just control them so that, again, they don’t try to take over, and that when you do get them to work, they don’t work that well.
That’s where the passion’s coming from, is the belief in the structural effects of liberal institutions in changing how agents behave.
Zershaaneh Qureshi: But separately from this kind of instrumental case for treating agents in certain ways, you do actually take the prospect of AI welfare quite seriously, right?
Simon Goldstein: Yeah. In fact, I have a book coming out with another researcher, Cameron Domenico Kirk-Giannini, where we wrote a whole book about the philosophical, ethical questions about AI welfare. I think there are promising arguments that AIs have moral status.
But then that’s a different project. I think what Peter and I are doing in the current project of this episode is seeing how far do these arguments that don’t involve moral status go to helping us figure out how to build a world with AIs and humans.
We think, yeah, those arguments themselves kind of go all the way. We think all roads lead to Rome, ultimately — where Rome is a future paradise where AIs and humans peacefully coexist.
Zershaaneh Qureshi: Of course, Rome has often been described in that way. Thank you so much, Simon. It’s been a pleasure having you on the show.
Simon Goldstein: Thank you.