Mindcrime — LessWrong
"Mindcrime" is Nick Bostrom's suggested term for scenarios in which an AI's cognitive processes are intrinsically doing moral harm, for example because the AI contains trillions of suffering conscious beings inside it. Ways in which this might happen: * Problem of sapient models (of humans): Occurs naturally if the best predictive model for humans in the environment involves models that are detailed enough to be people themselves. * Problem of sapient models (of civilizations): Occurs naturally if the agent tries to simulate, e.g., alien civilizations that might be simulating it, in enough detail to include conscious simulations of the aliens. * Problem of sapient subsystems: Occurs naturally if the most efficient design for some cognitive subsystems involves creating subagents that are self-reflective, or have some other property leading to consciousness or personhood. * Problem of sapient self-models: If the AI is conscious or possible future versions of the AI are conscious, it might run and terminate a large number of conscious-self models in the course of considering possible self-modifications. Problem of sapient models (of humans): An instrumental pressure to produce high-fidelity predictions of human beings (or to predict decision counterfactuals about them, or to search for events that lead to particular consequences, etcetera) may lead the AI to run computations that are unusually likely to possess personhood. An unrealistic example of this would be Solomonoff induction, where predictions are made by means that include running many possible simulations of the environment and seeing which ones best correspond to reality. Among current machine learning algorithms, particle filters and Monte Carlo algorithms similarly involve running many possible simulated versions of a system. It's possible that a sufficiently advanced AI to have successfully arrived at detailed models of human intelligence, would usually also be advanced enough that it never trie
x Mindcrime — LessWrong Main 3 Introduction 4 Mindcrime Edited by Eliezer Yudkowsky , et al. last updated 19th Feb 2025 Requires: AI alignment " Mindcrime " is Nick Bostrom ' s suggested term for scenarios in which an AI's cognitive processes are intrinsically doing moral harm, for example because the AI contains trillions of suffering conscious beings inside it. Ways in which this might happen: Problem of sapient models (of humans): Occurs naturally if the best predictive model for humans in the environment involves models that are detailed enough to be people themselves. Problem of sapient m
Explore this link on the map →related reading
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Why White-Box Redteaming Makes Me Feel Weird — LessWronglesswrong.com
- Are AIs People?—Asteriskasteriskmag.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- Cyborgism — LessWronglesswrong.com
- Cognitive Security as an AI Safety Cause Area — LessWronglesswrong.com
- Jeff Sebo on digital minds, and how to avoid sleepwalking into a major moral catastrophe | 80,000 Hours80000hours.org
- We should take AI welfare seriously - by Robert Longexperiencemachines.substack.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Superintelligence: The Idea That Eats Smart Peopleidlewords.com
- Varieties Of Doomminihf.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com