flâneur — a map of the web's best reading

Constitutional AI vs. RLHF vs. Deliberative Alignment — LessWrong

lesswrong.com · 2,753 words · saved by 1 readers

Outline: • 1. Quick review of RLHF, Constitutional AI, and Deliberative Alignment for a somewhat-technical audience, literature review of historical…

x Constitutional AI vs. RLHF vs. Deliberative Alignment — LessWrong AI Alignment Fieldbuilding AI Control AI Frontpage 26 Constitutional AI vs. RLHF vs. Deliberative Alignment by laudiacay 11th Apr 2026 10 min read 0 26 Outline: Quick review of RLHF, Constitutional AI, and Deliberative Alignment for a somewhat-technical audience, literature review of historical failure modes. Introduce "Persona-Emotion-Behavior space"- combining two recent interpretability papers to get a loose framework for talking about personality stability and current alignment techniques What's going on with alignment in

Explore this link on the map →

saved by

related reading