Fine-tuning GPT-2 from human preferences | OpenAI
We’ve fine-tuned the 774M parameter GPT‑2 language model using human feedback for various tasks, successfully matching the preferences of the external human labelers, though those preferences did not always match our own. Specifically, for summarization tasks the labelers preferred sentences copied wholesale from the input (we’d only asked them to ensure accuracy), so our models learned to copy. Summarization required 60k human labels; simpler tasks which continue text in various styles required only 5k. Our motivation is to move safety techniques closer to the general task of “machines talking to humans,” which we believe is key to extracting information about human values. We believe language is a key ingredient in making reinforcement learning practical and safe for real-world tasks. Previous work (opens in a new window) on learning models of human preferences has focused on simple simulated environments (Atari games or robotics tasks) which do not capture the complexity of langu
September 19, 2019 Publication Fine-tuning GPT‑2 from human preferences Read paper (opens in a new window) View code (opens in a new window) Loading… Share We’ve fine-tuned the 774M parameter GPT‑2 language model using human feedback for various tasks, successfully matching the preferences of the external human labelers, though those preferences did not always match our own. Specifically, for summarization tasks the labelers preferred sentences copied wholesale from the input (we’d only asked them to ensure accuracy), so our models learned to copy. Summarization required 60k human labels; simp
Explore this link on the map →saved by
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- gpt-4.pdfcdn.openai.com
- [2204.14146] Training Language Models with Language Feedbackarxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- rlhfbook.com/book.pdfrlhfbook.com
- Fine-Tuning Llama-2: Tailoring Models to Unique Applicationsanyscale.com
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- Recent Advances in Language Model Fine-tuningruder.io
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- [2507.06187] The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gainsarxiv.org