Reinforcement learning with imperceptible rewards — AI Alignment Forum
TLDR: We define a variant of reinforcement learning in which the reward is not perceived directly, but can be estimated at any given moment by some (…
x window.__lwSsrGql.inject("query PostsPageWrapper($documentId: String, $sequenceId: String) {\n post(input: {selector: {documentId: $documentId}}, allowNull: true) {\n result {\n ...PostsWithNavigation\n }\n }\n}\n\nfragment PostsMinimumInfo on Post {\n _id\n slug\n title\n draft\n shortform\n hideCommentKarma\n af\n userId\n coauthorUserIds\n rejected\n collabEditorDialogue\n}\n\nfragment PostsBase on Post {\n ...PostsMinimumInfo\n url\n postedAt\n sticky\n metaSticky\n stickyPriority\n status\n frontpageDate\n meta\n deletedDraft\n postCategory\n tagRelevance\n shareWithUsers\n sharingSetti
Explore this link on the map →related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward is not the optimization target — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Reinforcement learning - Wikipediaen.wikipedia.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- RUDDER - Reinforcement Learning with Delayed Rewards | rudderml-jku.github.io
- [2510.13651] What is the objective of reasoning with reinforcement learning?arxiv.org
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc
- [1906.09624] On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inferencearxiv.org