user_interactions.pdf
self-distillation.github.io · 8,336 words · saved by 1 readers
N/A
Aligning Language Models from User Interactions Aligning Language Models from User Interactions Thomas Kleine Buening1 Jonas Hübotter1 Barna Pásztor1 Idan Shenfeld2 Giorgia Ramponi3 Andreas Krause1 1 ETH Zurich 2 MIT 3 University of Zurich Abstract Multi-turn user interactions are among the most abundant data produced by language models, yet we lack effective methods to learn from them. While typically discarded, these interactions often contain useful information: follow-up user messages may…
saved by
related reading
- Checklists Are Better Than Reward Models For Aligning Language Modelsmachinelearning.apple.com
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com
- Simulating Users with State Alignment Beats Response Imitationhumanlm.stanford.edu
- LLMs Get Lost in Evolving User Intentarxiv.org
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- 2401.10020.pdfarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Training language models to follow instructions with human feedback.pdfproceedings.neurips.cc
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Self-Adapting Language Modelsarxiv.org
- Tiny Interaction Models - Rajan Agarwalrajan.sh