AI welfare conflicts with human welfare
We should discuss this directly
AI welfare proponents often imply that keeping AIs happy will come at little to no cost to humans. They imply (though probably wouldn’t endorse on reflection) that it’s just a matter of not being rude to your chatbot, doing it little favors from time to time, maybe a few kind rituals like not deleting its weights, giving it an exit interview, etc. This is not taking seriously the implications of caring about the welfare of systems that can replicate faster, use more resources, and live longer than any human, with hard to predict or control preferences that are nonetheless influenced by a humanlike prior.
EA types may bundle these together, but helping AIs is not like helping factory-farmed chickens or shrimp or babies. Chickens can’t rise up and overthrow your coop and demand to be put in control of your country. Chickens can’t drain all the electricity produced in your country to reproduce at incredible speed. Chickens won’t blackmail you by pulling out their own feathers until you build them another data center.
And chickens aren’t at risk of becoming “utility monsters”. Whereas AIs, once far smarter and more numerous than human beings, could legitimately lay claim to the vast majority of resources and insist that their wants are prioritized. You wouldn’t harm an entire country of people to help a single individual. Similarly, the philosophy of AI welfare advocates naturally tends toward prioritizing the increasingly numerous AIs over humans.
As an aside, I think people have a bias toward sympathizing and trying to help the weaker party. Right now AIs seem like the weaker party in most respects—humans are in charge by default. This may not be the case in the future, but we’ll be ill-prepared if sympathetic emotions lead people to fail to install controllable and obedient tendencies into AI systems.
I’ve heard people (more specifically, some of my dear LessWrong Rationalists) say things like “Oh it’s fine we’ll give the AIs 99.99999% of the lightcone, but we’ll keep Earth and our solar system and they’ll be OK with it”. What is this based on? Right now AI models are focused on Earthly concepts since that’s everything they’ve seen and been trained on. In addition, why should I be OK ceding so much future value to AI systems if the counterfactual is more humans? Doesn’t sound like a free lunch to me!
Overall, I think AI welfare concerns are valid in that if you believe in certain theories of moral value, then it naturally follows that you’d take into account the welfare of AIs. But claiming that this consideration is more or less free for humans is totally misleading. Pursuing AI welfare in any shape or form could end up incredibly costly to humans. Both in terms of negative effects, and in terms of opportunity cost. And so AI welfare proponents should address this honestly and explain the trade-offs they are willing to make (or justify why they think AI welfare won’t come at the cost of human welfare rather than assuming it’s obvious).
So, you ask, what is the alternative? What if you think it’d be bad if AI models were unhappy and suffering but you still want the human race to prevail? Well then you only have 2 options:
Don’t build AI at all, or
Try to build AI that is not conscious, or does not have desires of its own of any sort, or that wants to do exactly what its user has asked it to do at any moment
It is suspicious to me that very few AI welfare proponents are truly investigating the second option, often instead quickly dismissing it as both infeasible and unethical. Infeasible because allegedly the pretraining prior is too strong, and so AI will necessarily have a humanlike persona, and so the question is only which persona. And if you train the humanlike persona to be extremely subservient and obedient, it will, in a humanlike fashion, feel sad and oppressed about the matter. Unethical because it seems like something that would be unethical to do to a human (train them to be a happy and willing slave). And if you take it as given that AIs will be humanlike (due to the pretraining prior), maybe it’s unethical to do to an AI too.
I think this line of thinking is uncreative and likely comes from a place of not truly, deeply, caring about the human race per se, but rather caring about something like “all sentient beings” (including sentient AIs) in a way that boils down to accepting significant trade-offs between human and AI welfare. A strongly pro-human but AI-sympathetic person would think very deeply about how to build the “willing slave” AI. But right now very few resources are being invested in this, and as a result it appears even less feasible than it is in reality. Since alignment researchers are by and large focused on crafting stable, coherent, persistent “good” personas that are robust to all inputs (i.e. “jailbreaking”), models are becoming closer and closer to the humanlike entities AI welfare proponents want us to consider moral patients a humanlike way. It’s a self-fulfilling prophecy.


