flâneur — a map of the web's best reading

[2312.09390] Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

ar5iv.labs.arxiv.org · 24,729 words · saved by 1 readers

Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior—for example, to evaluate whether a model faithfully followed instructions or generated safe outputs. However, future superhuman models will behave in complex ways too difficult for humans to reliably evaluate; humans will only be able to weakly supervise superhuman models. We study an analogy to this problem: can weak model supervision elicit the full capabilities of a much stronger model? We test this using a range of pretrained language models in the GPT-4 family on natural language processing (NLP), chess, and reward modeling tasks. We find that when we naively finetune strong pretrained models on labels generated by a weak model, they consistently perform better than their weak supervisors, a phenomenon we call weak-to-strong generalization. However, we are still far from recovering the full capabilities of strong models with naive f

Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision Collin Burns &Pavel Izmailov ∗ &Jan Hendrik Kirchner ∗ &Bowen Baker ∗ &Leo Gao ∗ \AND Leopold Aschenbrenner ∗ &Yining Chen ∗ &Adrien Ecoffet ∗ &Manas Joglekar ∗ \AND Jan Leike &Ilya Sutskever &Jeff Wu ∗ \AND OpenAI Primary authors. This was a joint project of the Superalignment Generalization team. Correspondence to generalization@openai.com . Code is available at github.com/openai/weak-to-strong . Abstract Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the a

Explore this link on the map →

related reading