flâneur — a map of the web's best reading

diffuse.one

diffuse.one · 3,241 words · saved by 2 readers

Abstract: Designing a robust and correct reward function is difficult. Here I recount the long and iterative process of designing two reward functions. The first is for retrosynthesis of a target molecule and the second is generating a molecule with a specific number of atoms. These two are from a recent paper about building a chemical reasoning model. Reasoning models are compelling for science because of their strong accuracy and lower data requirements--they do not require worked out answers, only verifiers. Building these verifiers, the reward function, is challenging though. It requires knowing precisely what you want the model to be able to do, and that requires strong domain knowledge. Reinforcement learning is the process of training a reasoning model to get high scores on your reward function. Reinforcement learning is amazing, and perilous, because it reveals all the ways your reward function is misspecified and the models find ways to hack around this. Reward hacking is when

diffuse.one diffuse.one/ build_reward_functions designation: M1-000 author: andrew white status: complete topic: reward hacking in ether0 prepared date: June 20, 2025 updated date: June 22, 2025 abstract: Designing a robust and correct reward function is difficult. Here I recount the long and iterative process of designing two reward functions. The first is for retrosynthesis of a target molecule and the second is generating a molecule with a specific number of atoms. These two are from a recent paper about building a chemical reasoning model . building reward functions Reasoning models are co

Explore this link on the map →

saved by

related reading