Unsupervised Elicitation
Anthropic; 2Schmidt Sciences; 3Independent; 4Constellation; 5New York University; 6George Washington University We introduce a new unsupervised algorithm for eliciting skills from pretrained language models. This algorithm is competitive with training on human labels on common misconceptions (TruthfulQA), math (GSM8k-verification), and helpfulness reward modeling (Alpaca). Without supervision, we train a helpful chat assistant from the Haiku 3.5 base model that outperforms a similarly trained human-supervised baseline. 📄 Paper, 💻 Code A key problem in alignment research is how to align superhuman models whose behavior humans cannot reliably supervise. If we use today’s standard post-training approach to align models with human-specified behaviors (e.g., RLHF), we might train models to tell us what we want to hear even if it’s wrong, or do things that seem superficially good but are actually very different from what we intended. We introduce a new unsupervised algorithm to address thi
Unsupervised Elicitation Alignment Science Blog Unsupervised Elicitation Jiaxin Wen, Zachary Ankner, Arushi Somani, Peter Hase 2 , Samuel Marks, Jacob Goldman-Wetzler, Linda Petrini 3 , Henry Sleight 4 , Collin Burns, He He 5 , Shi Feng 6 , Ethan Perez, Jan Leike Anthropic; 2 Schmidt Sciences; 3 Independent; 4 Constellation; 5 New York University; 6 George Washington University tl;dr We introduce a new unsupervised algorithm for eliciting skills from pretrained language models. This algorithm is competitive with training on human labels on common misconceptions (TruthfulQA), math (GSM8k-verifi
Explore this link on the map →saved by
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- 3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitationalignment.anthropic.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2212.03827] Discovering Latent Knowledge in Language Models Without Supervisionarxiv.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org