flâneur — a map of the web's best reading

Trajectory Improvement and Reward Learning from Comparative Language Feedback

arxiv.org · 11,039 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

Trajectory Improvement and Reward Learning from Comparative Language Feedback Zhaojing Yang University of Southern California Miru Jun University of Southern California Jeremy Tien University of California, Berkeley Stuart J. Russell University of California, Berkeley Anca Dragan University of California, Berkeley Erdem Bıyık University of Southern California Abstract Learning from human feedback has gained traction in fields like robotics and natural language processing in recent years. While prior works mostly rely on human feedback in the form of comparisons, language is a preferable modali

Explore this link on the map →

related reading