flâneur — a map of the web's best reading

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

arxiv.org · 34,586 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Paul Gölz Cornell University paulgoelz@cornell.edu Nika Haghtalab UC Berkeley nika@berkeley.edu Kunhe Yang UC Berkeley kunheyang@berkeley.edu Abstract After pre-training, large language models are aligned with human preferences based on pairwise comparisons. State-of-the-art alignment methods (such as PPO-based RLHF and DPO) are built on the assumption of aligning with a single preference model, despite being deployed in settings

Explore this link on the map →

related reading