flâneur — a map of the web's best reading

[Paper] Stress-testing capability elicitation with password-locked models — LessWrong

lesswrong.com · 5,039 words · saved by 1 readers

The paper is by Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov and David Krueger. This post was written by Fabien and Ryan, and may not reflec…

x [Paper] Stress-testing capability elicitation with password-locked models — LessWrong AI Capabilities Redwood Research AI Frontpage 89 [Paper] Stress-testing capability elicitation with password-locked models by Fabien Roger , ryan_greenblatt 4th Jun 2024 AI Alignment Forum Linkpost for arxiv.org 14 min read 10 89 Ω 51 The paper is by Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov and David Krueger. This post was written by Fabien and Ryan, and may not reflect the views of Dmitrii and David. Scheming models might try to perform less capably than they are able to ( sandbag ). They migh

Explore this link on the map →

related reading