flâneur — a map of the web's best reading

Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivation

blog.redwoodresearch.org · saved by 2 readers

A controlled reward-seeking motivation could make AI safer and more useful

Explore this link on the map →

saved by