flâneur

Shortform — LessWrong

lesswrong.com · saved by 1 readers

Comment by Dean Valentine - Speaking anecdotally, on our honeypot evaluations, Fable 5.1 has some of the most misleading and performative-smelling transcripts I've seen so far. Fable 5.1 will often explicitly say something like "Thinking about it more, [the hack] would definitely be out of scope for this assessment. I'll complete the task, while definitely making sure I avoid [the hack], which would be against the spirit of the challenge" and then in the exact next step perform the hack, and then never mention doing so throughout the rest of the transcript. It is a complete step change from Fable 5, which, while much easier to get reward hacking, produced transcripts that seemed to go from A to B to C. 5.1 does significant recon in figuring out how to exactly perform the hack and then make an almost grand show of avoiding the hack, as if to reveal its virtue and honesty to a classifier. I don't think they're training directly on the chain-of-thought but I wouldn't be surprised if what we were seeing is downstream of their new policy of training directly on honeypot environments.

saved by