flâneur — a map of the web's best reading

Reward Is Not the Optimization Target

turntrout.com · 4,034 words · saved by 1 readers

RL doesn't train reward optimizers. Reward chisels cognition into agents. Worry less about safe objectives, more about good cognition.

Table of Contents Reward probably won’t be a deep RL agent’s primary optimization target The siren-like suggestiveness of the word “reward” When is reward the optimization target of the agent? Anticipated questions Dropping the old hypothesis Implications Citation Similar posts Appendix: The field of RL thinks reward is the optimization target Footnotes In this essay, I call an agent a “reward optimizer” if it not only gets lots of reward, but if it reliably makes choices like “reward but no task completion” (e.g. receiving reward without eating pizza) over “task completion but no reward” (e.g

Explore this link on the map →

related reading