2401.10020.pdf
arxiv.org · 7,977 words · saved by 1 readers
N/A
Self-Rewarding Language Models Weizhe Yuan1,2 Richard Yuanzhe Pang1,2 Kyunghyun Cho2 Xian Li1 Sainbayar Sukhbaatar1 Jing Xu1 Jason Weston1,2 1 2 Meta NYU arXiv:2401.10020v3 [cs.CL] 28 Mar 2025 Abstract…
saved by
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Fine-tune Llama 2 with DPOhuggingface.co
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Training language models to follow instructions with human feedback.pdfproceedings.neurips.cc
- RLHF | John Lambertjohnwlambert.github.io
- rlhfbook.com/book.pdfrlhfbook.com
- Checklists Are Better Than Reward Models For Aligning Language Modelsmachinelearning.apple.com
- [2305.14387] AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedbackarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimesarxiv.org