Reinforcement learning with imperceptible rewards — AI Alignment Forum
TLDR: We define a variant of reinforcement learning in which the reward is not perceived directly, but can be estimated at any given moment by some (…
x window.__lwSsrGql.inject("query PostsPageWrapper($documentId: String, $sequenceId: String) {\n post(input: {selector: {documentId: $documentId}}, allowNull: true) {\n result {\n ...PostsWithNavigation\n }\n }\n}\n\nfragment PostsMinimumInfo on Post {\n _id\n slug\n title\n draft\n shortform\n hideCommentKarma\n af\n userId\n coauthorUserIds\n rejected\n collabEditorDialogue\n}\n\nfragment PostsBase on Post {\n ...PostsMinimumInfo\n url\n postedAt\n sticky\n metaSticky\n stickyPriority\n status\n frontpageDate\n meta\n deletedDraft\n postCategory\n tagRelevance\n shareWithUsers\n sharingSetti
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward is not the optimization target — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- NeurIPS-2021-understanding-end-to-end-model-based-reinforcement-learning-methods-as-implicit-parameterization-Supplemental.pdflis.csail.mit.edu
- Reward Is Not the Optimization Targetturntrout.com
- RUDDER - Reinforcement Learning with Delayed Rewards | rudderml-jku.github.io
- Training a Misaligned Reward Seekeralignment.anthropic.com
- RLHF | John Lambertjohnwlambert.github.io
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reinforcement learning - Wikipediaen.wikipedia.org