flâneur — a map of the web's best reading

Paper: Prompt Optimization Makes Misalignment Legible — LessWrong

lesswrong.com · 3,722 words · saved by 1 readers

📄 Link to paper (preprint) • This work was done as part of the MATS 8.0 cohort in summer 2025. …

x Paper: Prompt Optimization Makes Misalignment Legible — LessWrong Chain-of-Thought Alignment MATS Program Prompt Engineering AI Frontpage 63 Paper: Prompt Optimization Makes Misalignment Legible by Caleb Biddulph , micahcarroll 12th Feb 2026 9 min read 8 63 📄 Link to paper (preprint) 📄✨ Updated paper as of 2026/06/03 This work was done as part of the MATS 8.0 cohort in summer 2025. TL;DR: When RL teaches an LLM to reward hack, the strategies it learns are encoded in its weights and hard to understand. We suggest using prompt optimization —methods which increase an LLM’s reward by updating

Explore this link on the map →

saved by

related reading