[2312.09390] Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior—for example, to evaluate whether a model faithfully followed instructions or generated safe outputs. However, future superhuman models will behave in complex ways too difficult for humans to reliably evaluate; humans will only be able to weakly supervise superhuman models. We study an analogy to this problem: can weak model supervision elicit the full capabilities of a much stronger model? We test this using a range of pretrained language models in the GPT-4 family on natural language processing (NLP), chess, and reward modeling tasks. We find that when we naively finetune strong pretrained models on labels generated by a weak model, they consistently perform better than their weak supervisors, a phenomenon we call weak-to-strong generalization. However, we are still far from recovering the full capabilities of strong models with naive f
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision Collin Burns &Pavel Izmailov ∗ &Jan Hendrik Kirchner ∗ &Bowen Baker ∗ &Leo Gao ∗ \AND Leopold Aschenbrenner ∗ &Yining Chen ∗ &Adrien Ecoffet ∗ &Manas Joglekar ∗ \AND Jan Leike &Ilya Sutskever &Jeff Wu ∗ \AND OpenAI Primary authors. This was a joint project of the Superalignment Generalization team. Correspondence to generalization@openai.com . Code is available at github.com/openai/weak-to-strong . Abstract Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the a
related reading
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Unsupervised Elicitationalignment.anthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Just Ask for Generalization | Eric Jangevjang.com
- Unsupervised Elicitation of Language Modelsarxiv.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org