flâneur — a map of the web's best reading

Learning from human preferences | OpenAI

openai.com · saved by 1 readers

We use cookies and similar technologies to deliver, maintain, improve our services and for security purposes. Check our Privacy Policy for details. Click 'Accept all' to let OpenAI and partners use cookies for these purposes. Click 'Reject non-essential' to say no to cookies, except those that are strictly necessary. June 13, 2017 One step towards building safe AI systems is to remove the need for humans to write goal functions, since using a simple proxy for a complex goal, or getting the complex goal a bit wrong, can lead to undesirable and even dangerous behavior. In collaboration with DeepMind’s safety team, we’ve developed an algorithm which can infer what humans want by being told which of two proposed behaviors is better. We present a learning algorithm that uses small amounts of human feedback to solve modern RL environments. Machine learning systems with human feedback have (opens in a new window) been (opens in a new window) explored (opens in a new window) before (opens i

We use cookies and similar technologies to deliver, maintain, improve our services and for security purposes. Check our Privacy Policy for details. Click 'Accept all' to let OpenAI and partners use cookies for these purposes. Click 'Reject non-essential' to say no to cookies, except those that are strictly necessary. June 13, 2017 One step towards building safe AI systems is to remove the need for humans to write goal functions, since using a simple proxy for a complex goal, or getting the complex goal a bit wrong, can lead to undesirable and even dangerous behavior. In collaboration with Deep

Explore this link on the map →