Constitutional AI: Harmlessness from AI Feedback
Abstract:As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and so we refer to the method as 'Constitutional AI'. The process involves both a supervised learning and a reinforcement learning phase. In the supervised phase we sample from an initial model, then generate self-critiques and revisions, and then finetune the original model on revised responses. In the RL phase, we sample from the finetuned model, use a model to evaluate which of the two samples is better, and then train a preference model from this dataset of AI preferences. We then train with RL using the preference model as the reward signal, i.e. we use 'RL from AI Feedback' (RLAIF). As a result we are able to train a harmless but non-evasive AI assistant that engages with harmful queries by explaining its objections to them. Both the SL and RL methods can leverage chain-of-thought style reasoning to improve the human-judged performance and transparency of AI decision making. These methods make it possible to control AI behavior more precisely and with far fewer human labels.
# link_1ormgktg7od.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20221219013858Z - Creator=LaTeX with hyperref - ModDate=D:20221219013858Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.14159265-2.6-1.40.21 (TeX Live 2020) kpathsea version 6.3.2 - Producer=pdfTeX-1.40.21 - Trapped=False ## Contents ### Page 1 Constitutional AI: Harmlessness from AI FeedbackYuntao Bai∗, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen,
Explore this link on the map →saved by
related reading
- [2212.08073] Constitutional AI: Harmlessness from AI Feedbackar5iv.labs.arxiv.org
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2310.13798] Specific versus General Principles for Constitutional AIarxiv.org
- Constitutional AI vs. RLHF vs. Deliberative Alignment — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2309.00267] RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedbackarxiv.org