Contra shard theory, in the context of the diamond maximizer problem — LessWrong
A bunch of my response to shard theory is a generalization of how niceness is unnatural. In a similar fashion, the other “shards” that the shard theory folk want to learn are unnatural too. That said, I'll spend a few extra words responding to the admirably-concrete diamond maximizer proposal that TurnTrout recently published, on the theory that briefly gesturing at my beliefs is better than saying nothing. I’ll be focusing on the diamond maximizer plan, though this criticism can be generalized and applied more broadly to shard theory. Finally, I'll note that the diamond maximization problem is not in fact the problem "build an AI that makes a little diamond", nor even "build an AI that probably makes a decent amount of diamond, while also spending lots of other resources on lots of other stuff" (although the latter is more progress than the former). The diamond maximization problem (as originally posed by MIRI folk) is a challenge of building an AI that definitely optimizes for a part
x Contra shard theory, in the context of the diamond maximizer problem — LessWrong 2022 MIRI Alignment Discussion Shard Theory AI Frontpage 107 Contra shard theory, in the context of the diamond maximizer problem by So8res 13th Oct 2022 AI Alignment Forum 3 min read 19 107 Ω 48 A bunch of my response to shard theory is a generalization of how niceness is unnatural . In a similar fashion, the other “shards” that the shard theory folk want to learn are unnatural too. That said, I'll spend a few extra words responding to the admirably-concrete diamond maximizer proposal that TurnTrout recently pu
Explore this link on the map →related reading
- The Shard Theory of Human Valuesturntrout.com
- Shard Theory in Nine Theses: a Distillation and Critical Appraisal — AI Alignment Forumalignmentforum.org
- The shard theory of human values — LessWronglesswrong.com
- Shard Theory — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- The shard theory of human values — AI Alignment Forumalignmentforum.org
- Understanding and avoiding value drift — LessWronglesswrong.com
- [April Fools'] Definitive confirmation of shard theory — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Have You Tried Thinking About It As Crystals? — LessWronglesswrong.com