Neuroscience of human social instincts: a sketch
In previous work, I have argued that we can divide the brain into a “Learning Subsystem” (cortex, striatum, etc.) housing randomly-initialized learning algorithms, and a “Steering Subsystem” (hypothalamus, brainstem, etc.) housing genetically-specified logic. Part of the Steering Subsystem is “human social instincts”—a suite of innate reactions and drives that are upstream of compassion, friendship, spite, norm-following, the sense of justice, and much more. The question addressed in this paper is: How do those human social instincts work? This problem is tricky because of a “symbol grounding problem”: Per above, I claim that our whole understanding of the world is built up by within-lifetime learning algorithms, and takes the form of a large unlabeled data structure. Certain activation states of this data structure—e.g., the activation state that represents someone insulting me—need to somehow trigger appropriate innate reactions. So there must be some way that the brain “grounds” these unlabeled learned concepts. How? In this article, I weave together many ideas supported by neuroscience research—visual heuristics in the superior colliculus, supervised learning in the amygdala, involuntary attention and learning rate modulation from the brainstem, and more—to sketch an answer. (26 pages, 21 figures)
Published May 27, 2026 | Version v3 1. Astera Institute Description In previous work, I have argued that we can divide the brain into a “Learning Subsystem” (cortex, striatum, etc.) housing randomly-initialized learning algorithms, and a “Steering Subsystem” (hypothalamus, brainstem, etc.) housing genetically-specified logic. Part of the Steering Subsystem is “human social instincts”—a suite of innate reactions and drives that are upstream of compassion, friendship, spite, norm-following, the sense of justice, and much more. The question addressed in this paper is: How do those human…
saved by
related reading
- Neuroscience of human social instincts: a sketch — LessWronglesswrong.com
- A Crash Course in the Neuroscience of Human Motivation — LessWronglesswrong.com
- The Shard Theory of Human Valuesturntrout.com
- My AGI safety research—2024 review, ’25 plans — AI Alignment Forumalignmentforum.org
- https://io0.github.ioio0.github.io
- My computational framework for the brain — LessWronglesswrong.com
- The shard theory of human values — LessWronglesswrong.com
- The Intense World Theory of Autism — LessWronglesswrong.com
- Reward Is Not Enough — LessWronglesswrong.com
- Longitudinal Science – Cross-disciplinary road-mapping and analysis of science and technology topicslongitudinal.blog
- So how does learning work?cs.mcgill.ca
- AGI safetysjbyrnes.com