Shard Theory - AI Alignment Forum
Shard theory is an alignment research program, about the relationship between training variables and learned values in trained Reinforcement Learning (RL) agents. It is thus an approach to progressively fleshing out a mechanistic account of human values, learned values in RL agents, and (to a lesser extent) the learned algorithms in ML generally. Shard theory's basic ontology of RL holds that shards are contextually activated, behavior-steering computations in neural networks (biological and artificial). The circuits that implement a shard that garners reinforcement are reinforced, meaning that that shard will be more likely to trigger again in the future, when given similar cognitive inputs. As an appreciable fraction of a neural network is composed of shards, large neural nets can possess quite intelligent constituent shards. These shards can be sophisticated enough to be well-modeled as playing negotiation games with each other, (potentially) explaining human psychological phenomena like akrasia and value changes from moral reflection. Shard theory also suggests an approach to explaining the shape of human values, and scheme for RL alignment.
x Shard Theory — AI Alignment Forum Shard Theory Edited by David Udell , ihatenumbersinusernames7 , et al. last updated 25th Jan 2026 Shard Theory is an alignment research program, about the relationship between training variables and learned values in trained Reinforcement Learning (RL) agents. It is thus an approach to progressively fleshing out a mechanistic account of human values , learned values in RL agents, and (to a lesser extent) the learned algorithms in ML generally. Shard theory's basic ontology of RL holds that shards are contextually activated, behavior-steering computations in
Explore this link on the map →related reading
- The Shard Theory of Human Valuesturntrout.com
- Shard Theory in Nine Theses: a Distillation and Critical Appraisal — AI Alignment Forumalignmentforum.org
- The shard theory of human values — LessWronglesswrong.com
- The shard theory of human values — AI Alignment Forumalignmentforum.org
- Contra shard theory, in the context of the diamond maximizer problem — LessWronglesswrong.com
- Understanding and avoiding value drift — LessWronglesswrong.com
- [April Fools'] Definitive confirmation of shard theory — LessWronglesswrong.com
- A Shard Theory of Everything - by Jason Hausenloyfirstscattering.com
- LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Have You Tried Thinking About It As Crystals? — LessWronglesswrong.com