NeurIPS Poster $\beta$-DPO: Direct Preference Optimization with Dynamic $\beta$
neurips.cc · 150 words · saved by 1 readers
The NeurIPS Logo above may be used on presentations. Right-click and choose download. It is a vector graphic and may be used at any scale.
Direct Preference Optimization (DPO) has emerged as a compelling approach for training Large Language Models (LLMs) to adhere to human preferences. However, the performance of DPO is sensitive to the fine-tuning of its trade-off parameter $\beta$, as well as to the quality of the preference data. We analyze the impact of $\beta$ and data quality on DPO, uncovering that optimal $\beta$ values vary with the informativeness of pairwise data. Addressing the limitations of static $\beta$ values, we introduce a novel framework that dynamically calibrates $\beta$ at the batch level, informed by…
saved by
related reading
- Fine-tune Llama 2 with DPOhuggingface.co
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Preference Tuning LLMs with Direct Preference Optimization Methodshuggingface.co
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferencesarxiv.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Is DPO Superior to PPO for LLM Alignment? A Comprehensive Studyarxiv.org
- RLHF | John Lambertjohnwlambert.github.io
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- 2401.10020.pdfarxiv.org