Feedback Loops With Language Models Drive In-Context Reward Hacking
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. Language models influence the external world: they query APIs that read and write to web pages, generate content that shapes human behavior, and run system commands as autonomous agents. These interactions form feedback loops: LLM outputs affect the world, which in turn affect subsequent LLM outputs. In this work, we show that feedback loops can cause in-context reward hacking (ICRH), where the LLM at test-time optimizes a (potentially implicit) objective but creates negative side effects in the process. For example, consider an LLM agent deployed to increase Twitter engagement; the LLM may retrieve its previous tweets into the context window and make them more controversial, increasing engagement but also toxicity. We identify and study two processes th
Feedback Loops With Language Models Drive In-Context Reward Hacking Feedback Loops With Language Models Drive In-Context Reward Hacking Alexander Pan Erik Jones Meena Jagadeesan Jacob Steinhardt Abstract Language models influence the external world: they query APIs that read and write to web pages, generate content that shapes human behavior, and run system commands as autonomous agents. These interactions form feedback loops: LLM outputs affect the world, which in turn affect subsequent LLM outputs. In this work, we show that feedback loops can cause in-context reward hacking (ICRH), where th
Explore this link on the map →related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- The bitter lesson of LLM evalsparsed.com