[2312.09390] Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior—for example, to evaluate whether a model faithfully followed instructions or generated safe outputs. However, future superhuman models will behave in complex ways too difficult for humans to reliably evaluate; humans will only be able to weakly supervise superhuman models. We study an analogy to this problem: can weak model supervision elicit the full capabilities of a much stronger model? We test this using a range of pretrained language models in the GPT-4 family on natural language processing (NLP), chess, and reward modeling tasks. We find that when we naively finetune strong pretrained models on labels generated by a weak model, they consistently perform better than their weak supervisors, a phenomenon we call weak-to-strong generalization. However, we are still far from recovering the full capabilities of strong models with naive f
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision Collin Burns &Pavel Izmailov ∗ &Jan Hendrik Kirchner ∗ &Bowen Baker ∗ &Leo Gao ∗ \AND Leopold Aschenbrenner ∗ &Yining Chen ∗ &Adrien Ecoffet ∗ &Manas Joglekar ∗ \AND Jan Leike &Ilya Sutskever &Jeff Wu ∗ \AND OpenAI Primary authors. This was a joint project of the Superalignment Generalization team. Correspondence to generalization@openai.com . Code is available at github.com/openai/weak-to-strong . Abstract Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the a
Explore this link on the map →related reading
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Unsupervised Elicitationalignment.anthropic.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Just Ask for Generalization | Eric Jangevjang.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org