flâneur — a map of the web's best reading

Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO

arxiv.org · 25,466 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO Ruizhe Shi Tsinghua University srz21@mails.tsinghua.edu.cn &Minhak Song ∗ * KAIST minhaksong@kaist.ac.kr Runlong Zhou University of Washington vectorzh@cs.washington.edu &Zihan Zhang University of Washington zihanz46@uw.edu Maryam Fazel University of Washington mfazel@uw.edu & Simon S. Du University of Washington ssdu@cs.washington.edu Equal contribution.Work done while Minhak Song was visiting the University of Washington. Abstract We present a fine-grained theoretical analysis of the performance gap between

Explore this link on the map →

related reading