✳flâneur — a map of the web's best reading
How well do models follow their constitutions? — LessWrong
lesswrong.com · 16,297 words · saved by 1 readers
This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. …
x How well do models follow their constitutions? — LessWrong MATS Program AI Frontpage 2026 Top Fifty: 14 % 100 How well do models follow their constitutions? by aryaj , Senthooran Rajamanoharan , Neel Nanda 12th Mar 2026 AI Alignment Forum 31 min read 5 100 Ω 33 This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. There's been a lot of buzz around Claude's 30K word constitution ("soul doc"), and unusual ways Anthropic is integrating it into training. If we can robustly train complex and nuanced values into a model, this would be a big deal for saf
Explore this link on the map →saved by
related reading
- (a) Anthropic constitution.arxiv.org
- Claude 4 System Cardwww-cdn.anthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- gpt-4.pdfcdn.openai.com
- Claude’s Constitution \ Anthropicanthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- Thoughts on Claude’s Constitution – Windows On Theorywindowsontheory.org
- Claude's Constitutional Structure - by Zvi Mowshowitzthezvi.substack.com
- confessions_paper.pdfcdn.openai.com
- Teaching Claude why \ Anthropicanthropic.com