✳flâneur — a map of the web's best reading
Where Do LLM Values Come From? — LessWrong
lesswrong.com · 5,708 words · saved by 1 readers
This work was done as part of the MATS 8.1 Program. …
x Where Do LLM Values Come From? — LessWrong Frontpage 17 Where Do LLM Values Come From? by lilysun004 , Arthur Conmy , Josh Engels 9th Jul 2026 21 min read 1 17 This work was done as part of the MATS 8.1 Program. 0: TL;DR LLMs learn "values": general considerations (e.g. "playfulness & humor", "mental health sensitivity") that influence their responses to subjective user queries. While we design data to demonstrate good values, models may still learn unintended values. We evaluate Olmo-3 ( Olmo et al. 2025 ) using a values eval ( Zhang et al., 2025 ) to show that values change over post-train
Explore this link on the map →related reading
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- The bitter lesson of LLM evalsparsed.com
- The persona selection model — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org