On Anthropic's Sleeper Agents Paper - by Zvi Mowshowitz
Scott Alexander also covers this, offering an excellent high level explanation, of both the result and the arguments about whether it is meaningful. You could start with his write-up to get the gist, then return here if you still want more details, or you can read here knowing that everything he discusses is covered below. There was one good comment, pointing out some of the ways deceptive behavior could come to pass, but most people got distracted by the ‘grue’ analogy. Right up front before proceeding, to avoid a key misunderstanding: I want to emphasize that in this paper, the deception was introduced intentionally. The paper deals with attempts to remove it. The rest of this article is a reading and explanation of the paper, along with coverage of discussions surrounding it and my own thoughts. Paper Abstract: Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives wh
The recent paper from Anthropic is getting unusually high praise, much of it I think deserved. The title is: Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. Scott Alexander also covers this, offering an excellent high level explanation, of both the result and the arguments about whether it is meaningful. You could start with his write-up to get the gist, then return here if you still want more details, or you can read here knowing that everything he discusses is covered below. There was one good comment, pointing out some of the ways deceptive behavior could…
related reading
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Deep Deceptiveness — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- 2401.05566.pdfarxiv.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Self-Other Overlap: A Neglected Approach to AI Alignment — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Natural Deception with RL - Rajan Agarwalrajan.sh
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org