Human exploration hacking | Rohan's Rambles
One potential failure mode of automating AI safety research (i.e., using existing models to align and monitor subsequent ones) is that the models we try to use for that purpose are misaligned, and actually end up sabotaging our work. These sabotage strategies could be insidious – for example, the model doing a sloppy job of monitoring even if it’s capable of performing better – or more brazen – e.g., secretly inserting vulnerabilities into the code it writes for future models to exploit. Perhaps we try to remove these behaviours and train the models to be genuinely helpful by applying reinforcement learning (RL) that rewards effective monitoring and useful safety research. But this countermeasure could itself fail: a misaligned model might fight back by “exploration hacking” – deliberately avoiding exploring the actions it thinks we’ll reward. The RL training would be ineffective as a result (since there wouldn’t be any high-value actions identified for it to reinforce), allowing the m
Human exploration hacking | Rohan's Rambles Human exploration hacking First published: 05 November 2025 With many thanks to Emil for helpful comments and discussion. One potential failure mode of automating AI safety research (i.e., using existing models to align and monitor subsequent ones) is that the models we try to use for that purpose are misaligned, and actually end up sabotaging our work. These sabotage strategies could be insidious - for example, the model doing a sloppy job of monitoring even if it's capable of performing better - or more brazen - e.g., secretly inserting vulnerabili
Explore this link on the map →related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hackingblog.redwoodresearch.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- How can we solve diffuse threats like research sabotage with AI control? — LessWronglesswrong.com
- Ten people on the inside — LessWronglesswrong.com