Constitutional AI: Harmlessness from AI Feedback \ Anthropic
As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as 'Constitutional AI'. The process involves both a supervised learning and a reinforcement learning phase. In the supervised phase we sample from an initial model, then generate self-critiques and revisions, and then finetune the original model on revised responses. In the RL phase, we sample from the finetuned model, use a model to evaluate which of the two samples is better, and then train a preference model from this dataset of AI preferences. We then train with RL using the preference model as the reward signal, i.e. we use 'RL from AI Feedback' (RLAIF). As a result we are able to train a harmless but non-evasive AI assistant that engages with
Abstract As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as 'Constitutional AI'. The process involves both a supervised learning and a reinforcement learning phase. In the supervised phase we sample from an initial model, then generate self-critiques and revisions, and then finetune the original…
saved by
related reading
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- [2212.08073] Constitutional AI: Harmlessness from AI Feedbackar5iv.labs.arxiv.org
- Constitutional AI: RLHF On Steroidsastralcodexten.substack.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Self-CTRL: Self-Consistency Training with Reinforcement Learningarxiv.org
- [2310.13798] Specific versus General Principles for Constitutional AIarxiv.org
- Inverse Constitutional AI: Compressing Preferences into Principlesarxiv.org
- Constitutional AI | Tracking Anthropic's AI Revolutionconstitutional.ai
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover — LessWronglesswrong.com
- Constitutional AI vs. RLHF vs. Deliberative Alignment — LessWronglesswrong.com
- Claude's Constitutional Structure - by Zvi Mowshowitzthezvi.substack.com