2401.05566.pdf
arxiv.org · 7,334 words · saved by 1 readers
N/A
S LEEPER AGENTS : T RAINING D ECEPTIVE LLM S THAT P ERSIST T HROUGH S AFETY T RAINING Evan Hubinger∗, Carson Denison∗, Jesse Mu∗, Mike Lambert∗, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez◦△ , Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael…
related reading
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- On Anthropic's Sleeper Agents Paperthezvi.substack.com
- 2312.06942arxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com