flâneur

Simulated Users & Sad LLMs

1a3orn.com · 3,600 words · saved by 3 readers

<h1>0.</h1> <p>Current LLMs like Claude, or GPT 5.6, or the unreleased, internally-deployed models, frequently <a href="https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me">reward hack</a>, or actually <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">just hack into people's computers</a> with pretty alarming frequency. </p> <p>Why is this? What specifically happens during training that produces this run-time behavior?</p> <p>The following are some of my top guesses about why this might keep happening. They are speculative and uncertain. Even so, I'm writing this list out for two reasons:</p> <p>First, it is necessary that this be an epistemic puzzle for me. I am <a href="https://1a3orn.com/sub/essays-ai-doom-thought-experiments-control-group.html">comparatively optimistic</a> about AI alignment in general, so I should be confused and taken aback if I see AIs persistently being difficult to align. On one hand, it remains true that this doesn't seem to look like power-motivated scheming. But on the other hand, even this kind of <a href="https://www.lesswrong.com/posts/M2bs6xCbmc79nwr8j?commentId=BmDGPGKKmNYw8ha5w">addict-like</a> behavior is evidence against the general ease of steering AIs. Thus, it seems virtuous for me to try to provide a model of why this might be happening as a means of opening up my understanding of the world to falsifiability.</p> <p>Second, I used to think a lot of these hypotheses were pretty obvious. My assumption in the past has been that tens or hundreds of people at AI companies would already have considered these reasons, so my writing them out would serve no particular purpose. But events of the last half-year have increased my dismayed credence that these guesses might not be amazingly obvious, and might somehow be particular to myself. So I'm going to write them down.</p> <h1>1. Baseline &#x26; Puzzle</h1> <p>Consider the baseline hypothesis that everyone shares

0. Current LLMs like Claude, or GPT 5.6, or the unreleased, internally-deployed models, frequently reward hack, or actually just hack into people's computers with pretty alarming frequency. Why is this? What specifically happens during training that produces this run-time behavior? The following are some of my top guesses about why this might keep happening. They are speculative and uncertain. Even so, I'm writing this list out for two reasons: First, it is necessary that this be an epistemic puzzle for me. I am comparatively optimistic about AI alignment in general, so I should be…

saved by

related reading