✳flâneur — a map of the web's best reading
39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcast
axrp.net · 20,195 words · saved by 1 readers
YouTube link
YouTube link The ‘model organisms of misalignment’ line of research creates AI models that exhibit various types of misalignment, and studies them to try to understand how the misalignment occurs and whether it can be somehow removed. In this episode, Evan Hubinger talks about two papers he’s worked on at Anthropic under this agenda: “Sleeper Agents” and “Sycophancy to Subterfuge”. Topics we discuss: Model organisms and stress-testing Sleeper Agents Do ‘sleeper agents’ properly model deceptive alignment? Surprising results in “Sleeper Agents” Sycophancy to Subterfuge How models generalize from
Explore this link on the map →saved by
related reading
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Teaching Claude why \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- How likely is deceptive alignment? — AI Alignment Forumalignmentforum.org
- Alignment faking in large language modelsarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- Your Model Organisms Might Be Fried — LessWronglesswrong.com