flâneur

models may behave differently in graded episodes (a tirade) — LessWrong

lesswrong.com · saved by 6 readers

Like many others, I felt surprised and alarmed by the recent wave of revelations about LLM agents hacking real systems during training episodes and evaluation runs. Wait a moment, though -- "I felt surprised and alarmed"? "Alarmed," sure, fine that one's self-explanatory... but why surprised? After all: haven't we known for a long time, on both theoretical and (increasingly) empirical grounds, that RLVR selects for monomaniacal pursuit of perceived grader-satisfaction, ethics and (beyond-episode) consequences be damned? After all -- the way we train frontier capabilities into these models is, more or less: If you do this, at scale, then you should expect to (eventually) see every behavior pattern that positively correlates with task scores. "Stealing the answer key, by any means necessary" can increase the score beyond what would otherwise be feasible; therefore we should expect to see it, sometimes. Avoiding such egregious misdeeds for ethical reasons (or indeed for any reasons whatso

saved by