FleetingBits.io
Here, I collect questions that I find interesting. If you write a response to one of them and you let me know and I think it's a banger then I will link it on my website. Here is a list of projects that I think people should do. I'm starting a new series of posts about interesting journal articles I've read, where each post will break down the paper's content, outline potential follow-up experiments, highlight connected research, and explore how the ideas fit into the larger AI landscape. This first post is about Monitoring Frontier Models for Misbehavior and the Risks of Promoting Obfuscation. The paper describes experiments conducted by OpenAI that tested whether monitors can prevent reasoning models from developing reward hacking strategies. Modern reasoning models are trained using a mix of perfect and imperfect verifiers. The problem with imperfect verifiers is that models can produce states that the verifier will judge as correct solutions but which are not correct. This can caus