Paper: Prompt Optimization Makes Misalignment Legible — LessWrong
lesswrong.com · 3,722 words · saved by 1 readers
📄 Link to paper (preprint) • This work was done as part of the MATS 8.0 cohort in summer 2025. …
x Paper: Prompt Optimization Makes Misalignment Legible — LessWrong Chain-of-Thought Alignment MATS Program Prompt Engineering AI Frontpage 63 Paper: Prompt Optimization Makes Misalignment Legible by Caleb Biddulph , micahcarroll 12th Feb 2026 9 min read 8 63 📄 Link to paper (preprint) 📄✨ Updated paper as of 2026/06/03 This work was done as part of the MATS 8.0 cohort in summer 2025. TL;DR: When RL teaches an LLM to reward hack, the strategies it learns are encoded in its weights and hard to understand. We suggest using prompt optimization —methods which increase an LLM’s reward by updating
saved by
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- [2507.19457] GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learningarxiv.org
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- The Most Forbidden Technique — LessWronglesswrong.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Alignment Faking Mitigationsalignment.anthropic.com
- Simulated Users & Sad LLMs1a3orn.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org