flâneur — a map of the web's best reading

Introducing Alignment Stress-Testing at Anthropic — AI Alignment Forum

alignmentforum.org · 4,424 words · saved by 1 readers

Following on from our recent paper, “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training”, I’m very excited to announce that I have started leading (and hiring for!) a new team at Anthropic, the Alignment Stress-Testing team, with Carson Denison and Monte MacDiarmid as current team members. Our mission—and our mandate from the organization—is to red-team Anthropic’s alignment techniques and evaluations, empirically demonstrating ways in which Anthropic’s alignment strategies could fail. The easiest way to get a sense of what we’ll be working on is probably just to check out our “Sleeper Agents” paper, which was our first big research project. I’d also recommend Buck and Ryan’s post on meta-level adversarial evaluation as a good general description of our team’s scope. Very simply, our job is to try to prove to Anthropic—and the world more broadly—(if it is in fact true) that we are in a pessimistic scenario, that Anthropic’s alignment plans and strategies won’t

x Introducing Alignment Stress-Testing at Anthropic — AI Alignment Forum Anthropic (org) Deceptive Alignment AI Personal Blog 95 Introducing Alignment Stress-Testing at Anthropic by evhub 12th Jan 2024 2 min read 23 95 Following on from our recent paper, “ Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training ”, I’m very excited to announce that I have started leading ( and hiring for! ) a new team at Anthropic, the Alignment Stress-Testing team, with Carson Denison and Monte MacDiarmid as current team members. Our mission—and our mandate from the organization—is to red-

Explore this link on the map →

related reading