✳flâneur — a map of the web's best reading
Paper: Prompt Optimization Makes Misalignment Legible — LessWrong
lesswrong.com · 3,722 words · saved by 1 readers
📄 Link to paper (preprint) • This work was done as part of the MATS 8.0 cohort in summer 2025. …
x Paper: Prompt Optimization Makes Misalignment Legible — LessWrong Chain-of-Thought Alignment MATS Program Prompt Engineering AI Frontpage 63 Paper: Prompt Optimization Makes Misalignment Legible by Caleb Biddulph , micahcarroll 12th Feb 2026 9 min read 8 63 📄 Link to paper (preprint) 📄✨ Updated paper as of 2026/06/03 This work was done as part of the MATS 8.0 cohort in summer 2025. TL;DR: When RL teaches an LLM to reward hack, the strategies it learns are encoded in its weights and hard to understand. We suggest using prompt optimization —methods which increase an LLM’s reward by updating
Explore this link on the map →saved by
related reading
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- [2507.19457] GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learningarxiv.org
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- The Most Forbidden Technique — LessWronglesswrong.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- How hard is it to inoculate against misalignment generalization? — LessWronglesswrong.com
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai