Anthropic \ Collective Constitutional AI: Aligning a Language Model…
Anthropic and the Collective Intelligence Project recently ran a public input process involving ~1,000 Americans to draft a constitution for an AI system. We did this to explore how democratic processes can influence AI development. In our experiment, we discovered areas where people both agreed with our in-house constitution, and areas where they had different preferences. In this post, we share the resulting publicly sourced constitution, as well as what happened when we trained a new AI system against it using Constitutional AI. Constitutional AI (CAI) is an Anthropic-developed method for aligning general purpose language models to abide by high-level normative principles written into a constitution. Anthropic’s language model Claude currently relies on a constitution curated by Anthropic employees. This constitution takes inspiration from outside sources like the United Nations Universal Declaration of Human Rights, as well as our own firsthand experience interacting with language
Policy Societal Impacts Collective Constitutional AI: Aligning a Language Model with Public Input Oct 17, 2023 Anthropic and the Collective Intelligence Project recently ran a public input process involving ~1,000 Americans to draft a constitution for an AI system. We did this to explore how democratic processes can influence AI development. In our experiment, we discovered areas where people both agreed with our in-house constitution , and areas where they had different preferences. In this post, we share the resulting publicly sourced constitution, as well as what happened when we trained a
Explore this link on the map →related reading
- Collective Constitutional AI: Aligning a Language Model with Public Input \ Anthropicanthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- Claude's Constitutional Structure - by Zvi Mowshowitzthezvi.substack.com
- Constitutional AI | Tracking Anthropic's AI Revolutionconstitutional.ai
- Thoughts on Claude’s Constitution – Windows On Theorywindowsontheory.org
- Claude’s Constitution \ Anthropicanthropic.com
- [2310.13798] Specific versus General Principles for Constitutional AIarxiv.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Open Problems With Claude’s Constitution — LessWronglesswrong.com
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- Constitutional AI vs. RLHF vs. Deliberative Alignment — LessWronglesswrong.com
- How well do models follow their constitutions? — LessWronglesswrong.com