Deliberative Alignment, And The Spec - by Scott Alexander
astralcodexten.com · 2,181 words · saved by 1 readers
...
In the past day, Zvi has written about deliberative alignment, and OpenAI has updated their spec. This article was written before either of these and doesn’t account for them, sorry. OpenAI has bad luck with its alignment teams. The first team quit en masse to found Anthropic, now a major competitor. The second team quit en masse to protest the company reneging on safety commitments. The third died in a tragic plane crash. The fourth got washed away in a flood. The fifth through eighth were all slain by various types of wild beast. But the ninth team is still there and doing good work.…
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Constitutional AI vs. RLHF vs. Deliberative Alignment — LessWronglesswrong.com
- The Artificiality of Alignmentjoinreboot.org
- Teaching Claude why \ Anthropicanthropic.com
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com
- Thoughts on Claude’s Constitution – Windows On Theorywindowsontheory.org
- Prologue to Terrified Comments on Claude’s Constitution | An Algorithmic Lucidityzackmdavis.net
- Model Spec Midtraining: Improving How Alignment Training Generalizesalignment.anthropic.com
- [2605.24229] How Well Do Models Follow Their Constitutions?arxiv.org