flâneur — a map of the web's best reading

Unsupervised Elicitation

alignment.anthropic.com · 819 words · saved by 2 readers

Anthropic; 2Schmidt Sciences; 3Independent; 4Constellation; 5New York University; 6George Washington University We introduce a new unsupervised algorithm for eliciting skills from pretrained language models. This algorithm is competitive with training on human labels on common misconceptions (TruthfulQA), math (GSM8k-verification), and helpfulness reward modeling (Alpaca). Without supervision, we train a helpful chat assistant from the Haiku 3.5 base model that outperforms a similarly trained human-supervised baseline. 📄 Paper, 💻 Code A key problem in alignment research is how to align superhuman models whose behavior humans cannot reliably supervise. If we use today’s standard post-training approach to align models with human-specified behaviors (e.g., RLHF), we might train models to tell us what we want to hear even if it’s wrong, or do things that seem superficially good but are actually very different from what we intended. We introduce a new unsupervised algorithm to address thi

Unsupervised Elicitation Alignment Science Blog Unsupervised Elicitation Jiaxin Wen, Zachary Ankner, Arushi Somani, Peter Hase 2 , Samuel Marks, Jacob Goldman-Wetzler, Linda Petrini 3 , Henry Sleight 4 , Collin Burns, He He 5 , Shi Feng 6 , Ethan Perez, Jan Leike Anthropic; 2 Schmidt Sciences; 3 Independent; 4 Constellation; 5 New York University; 6 George Washington University tl;dr We introduce a new unsupervised algorithm for eliciting skills from pretrained language models. This algorithm is competitive with training on human labels on common misconceptions (TruthfulQA), math (GSM8k-verifi

Explore this link on the map →

saved by

related reading