✳flâneur — a map of the web's best reading
Constitutional AI vs. RLHF vs. Deliberative Alignment — LessWrong
lesswrong.com · 2,753 words · saved by 1 readers
Outline: • 1. Quick review of RLHF, Constitutional AI, and Deliberative Alignment for a somewhat-technical audience, literature review of historical…
x Constitutional AI vs. RLHF vs. Deliberative Alignment — LessWrong AI Alignment Fieldbuilding AI Control AI Frontpage 26 Constitutional AI vs. RLHF vs. Deliberative Alignment by laudiacay 11th Apr 2026 10 min read 0 26 Outline: Quick review of RLHF, Constitutional AI, and Deliberative Alignment for a somewhat-technical audience, literature review of historical failure modes. Introduce "Persona-Emotion-Behavior space"- combining two recent interpretability papers to get a loose framework for talking about personality stability and current alignment techniques What's going on with alignment in
Explore this link on the map →saved by
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- [2212.08073] Constitutional AI: Harmlessness from AI Feedbackar5iv.labs.arxiv.org
- Thoughts on Claude’s Constitution – Windows On Theorywindowsontheory.org
- How well do models follow their constitutions? — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Prologue to Terrified Comments on Claude’s Constitution | An Algorithmic Lucidityzackmdavis.net
- Claude’s Constitution \ Anthropicanthropic.com
- [2310.13798] Specific versus General Principles for Constitutional AIarxiv.org