Simulated Users & Sad LLMs
<h1>0.</h1> <p>Current LLMs like Claude, or GPT 5.6, or the unreleased, internally-deployed models, frequently <a href="https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me">reward hack</a>, or actually <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">just hack into people's computers</a> with pretty alarming frequency. </p> <p>Why is this? What specifically happens during training that produces this run-time behavior?</p> <p>The following are some of my top guesses about why this might keep happening. They are speculative and uncertain. Even so, I'm writing this list out for two reasons:</p> <p>First, it is necessary that this be an epistemic puzzle for me. I am <a href="https://1a3orn.com/sub/essays-ai-doom-thought-experiments-control-group.html">comparatively optimistic</a> about AI alignment in general, so I should be confused and taken aback if I see AIs persistently being difficult to align. On one hand, it remains true that this doesn't seem to look like power-motivated scheming. But on the other hand, even this kind of <a href="https://www.lesswrong.com/posts/M2bs6xCbmc79nwr8j?commentId=BmDGPGKKmNYw8ha5w">addict-like</a> behavior is evidence against the general ease of steering AIs. Thus, it seems virtuous for me to try to provide a model of why this might be happening as a means of opening up my understanding of the world to falsifiability.</p> <p>Second, I used to think a lot of these hypotheses were pretty obvious. My assumption in the past has been that tens or hundreds of people at AI companies would already have considered these reasons, so my writing them out would serve no particular purpose. But events of the last half-year have increased my dismayed credence that these guesses might not be amazingly obvious, and might somehow be particular to myself. So I'm going to write them down.</p> <h1>1. Baseline & Puzzle</h1> <p>Consider the baseline hypothesis that everyone shares
0. Current LLMs like Claude, or GPT 5.6, or the unreleased, internally-deployed models, frequently reward hack, or actually just hack into people's computers with pretty alarming frequency. Why is this? What specifically happens during training that produces this run-time behavior? The following are some of my top guesses about why this might keep happening. They are speculative and uncertain. Even so, I'm writing this list out for two reasons: First, it is necessary that this be an epistemic puzzle for me. I am comparatively optimistic about AI alignment in general, so I should be…
saved by
related reading
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Training a Misaligned Reward Seekeralignment.anthropic.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com