[Paper] Stress-testing capability elicitation with password-locked models — LessWrong
The paper is by Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov and David Krueger. This post was written by Fabien and Ryan, and may not reflec…
x [Paper] Stress-testing capability elicitation with password-locked models — LessWrong AI Capabilities Redwood Research AI Frontpage 89 [Paper] Stress-testing capability elicitation with password-locked models by Fabien Roger , ryan_greenblatt 4th Jun 2024 AI Alignment Forum Linkpost for arxiv.org 14 min read 10 89 Ω 51 The paper is by Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov and David Krueger. This post was written by Fabien and Ryan, and may not reflect the views of Dmitrii and David. Scheming models might try to perform less capably than they are able to ( sandbag ). They migh
related reading
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- AI in 2025: gestalt — LessWronglesswrong.com
- A structured protocol for elicitation experiments | AISI Workaisi.gov.uk
- Memorizing weak examples can elicit strong behavior out of password-locked models — LessWronglesswrong.com
- Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hackingblog.redwoodresearch.org
- Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hackingredwoodresearch.substack.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Auditing Games for Sandbagging [paper] — LessWronglesswrong.com