flâneur — a map of the web's best reading

Claude’s Constitution \ Anthropic

anthropic.com · 3,193 words · saved by 1 readers

How does a language model decide which questions it will engage with and which it deems inappropriate? Why will it encourage some actions and discourage others? What “values” might a language model have? These are all questions people grapple with. Our recently published research on “Constitutional AI” provides one answer by giving language models explicit values determined by a constitution, rather than values determined implicitly via large-scale human feedback. This isn’t a perfect approach, but it does make the values of the AI system easier to understand and easier to adjust as needed. Since launching Claude, our AI assistant trained with Constitutional AI, we've heard more questions about Constitutional AI and how it contributes to making Claude safer and more helpful. In this post, we explain what constitutional AI is, what the values in Claude’s constitution are, and how we chose them. If you just want to skip to the principles, scroll down to the last section which is entit

Announcements Claude’s Constitution May 9, 2023 Read the new constitution Update, Jan 21, 2026: We've published a new version of Claude's constitution, which you can find at the button above. How does a language model decide which questions it will engage with and which it deems inappropriate? Why will it encourage some actions and discourage others? What “values” might a language model have? These are all questions people grapple with. Our recently published research on “Constitutional AI” provides one answer by giving language models explicit values determined by a constitution, rather than

Explore this link on the map →

related reading