flâneur — a map of the web's best reading

Deep Reinforcement Learning from Human Preferences

proceedings.neurips.cc · 4,998 words · saved by 1 readers

N/A

# link_1v5y0rmw15d.pdf ## Metadata - PDFFormatVersion=1.3 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Subject=Neural Information Processing Systems http://nips.cc/ - Custom.Publisher=Curran Associates, Inc. - Custom.Language=en-US - Custom.Created=2017 - Custom.EventType=Poster - Custom.Description-Abstract=For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms o

Explore this link on the map →

saved by

related reading