Thousand-dimensional structure — Resolution
resolution.org · 2,858 words · saved by 3 readers
Finding and controlling the low-dimensional persona structure in models — from emergent misalignment and subliminal learning through to superintelligence.
Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional structure in models that emerges in pretraining and flows through post-training to superintelligence. The hope is to expand and systematize phenomena such as emergent misalignment, subliminal learning, and other empirical persona research, then intervene on this structure without accidentally hiding undesirable behavior elsewhere. Glimmers of low-dimensional structure Our understanding of AI training and alignment as a field is very poor. If…
saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com