Claude’s Constitution \ Anthropic
How does a language model decide which questions it will engage with and which it deems inappropriate? Why will it encourage some actions and discourage others? What “values” might a language model have? These are all questions people grapple with. Our recently published research on “Constitutional AI” provides one answer by giving language models explicit values determined by a constitution, rather than values determined implicitly via large-scale human feedback. This isn’t a perfect approach, but it does make the values of the AI system easier to understand and easier to adjust as needed. Since launching Claude, our AI assistant trained with Constitutional AI, we've heard more questions about Constitutional AI and how it contributes to making Claude safer and more helpful. In this post, we explain what constitutional AI is, what the values in Claude’s constitution are, and how we chose them. If you just want to skip to the principles, scroll down to the last section which is entit
Announcements Claude’s Constitution May 9, 2023 Read the new constitution Update, Jan 21, 2026: We've published a new version of Claude's constitution, which you can find at the button above. How does a language model decide which questions it will engage with and which it deems inappropriate? Why will it encourage some actions and discourage others? What “values” might a language model have? These are all questions people grapple with. Our recently published research on “Constitutional AI” provides one answer by giving language models explicit values determined by a constitution, rather than
Explore this link on the map →related reading
- Claude's Constitutional Structure - by Zvi Mowshowitzthezvi.substack.com
- Claude’s Constitution \ Anthropicanthropic.com
- Thoughts on Claude’s Constitution – Windows On Theorywindowsontheory.org
- Open Problems With Claude’s Constitution — LessWronglesswrong.com
- Claude's new constitution \ Anthropicanthropic.com
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- [2310.13798] Specific versus General Principles for Constitutional AIarxiv.org
- Constitutional AI | Tracking Anthropic's AI Revolutionconstitutional.ai
- Collective Constitutional AI: Aligning a Language Model with Public Input \ Anthropicanthropic.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- Values in the wild: Discovering and analyzing values in real-world language model interactions \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com