flâneur — a map of the web's best reading

Human exploration hacking | Rohan's Rambles

selvaradov.net · 850 words · saved by 1 readers

One potential failure mode of automating AI safety research (i.e., using existing models to align and monitor subsequent ones) is that the models we try to use for that purpose are misaligned, and actually end up sabotaging our work. These sabotage strategies could be insidious – for example, the model doing a sloppy job of monitoring even if it’s capable of performing better – or more brazen – e.g., secretly inserting vulnerabilities into the code it writes for future models to exploit. Perhaps we try to remove these behaviours and train the models to be genuinely helpful by applying reinforcement learning (RL) that rewards effective monitoring and useful safety research. But this countermeasure could itself fail: a misaligned model might fight back by “exploration hacking” – deliberately avoiding exploring the actions it thinks we’ll reward. The RL training would be ineffective as a result (since there wouldn’t be any high-value actions identified for it to reinforce), allowing the m

Human exploration hacking | Rohan's Rambles Human exploration hacking First published: 05 November 2025 With many thanks to Emil for helpful comments and discussion. One potential failure mode of automating AI safety research (i.e., using existing models to align and monitor subsequent ones) is that the models we try to use for that purpose are misaligned, and actually end up sabotaging our work. These sabotage strategies could be insidious - for example, the model doing a sloppy job of monitoring even if it's capable of performing better - or more brazen - e.g., secretly inserting vulnerabili

Explore this link on the map →

related reading