Naturally Occurring Feedback is Common, Extractable and Useful
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Naturally Occurring Feedback is Common, Extractable and Useful Shachar Don-Yehiya 1 Leshem Choshen 2,3 Omri Abend 1 1 The Hebrew University of Jerusalem, 2 MIT, 3 MIT-IBM Watson AI Lab {first.last}@mail.huji.ac.il Abstract Human feedback data is a critical component in developing language models. However, collecting this feedback is costly and ultimately not scalable. Inspired by the way human interlocutors provide spontaneous unsolicited feedback to each other, we propose to extract feedback that users naturally include when interacting with chat models. We manually annotated conversations to
Explore this link on the map →related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- [2204.14146] Training Language Models with Language Feedbackarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- 2405.01470arxiv.org
- rlhfbook.com/book.pdfrlhfbook.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Unsupervised Elicitationalignment.anthropic.com
- [2606.06614] Re-Centering Humans in LLM Personalizationarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org