[Paper] Stress-testing capability elicitation with password-locked models — LessWrong
The paper is by Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov and David Krueger. This post was written by Fabien and Ryan, and may not reflec…
x [Paper] Stress-testing capability elicitation with password-locked models — LessWrong AI Capabilities Redwood Research AI Frontpage 89 [Paper] Stress-testing capability elicitation with password-locked models by Fabien Roger , ryan_greenblatt 4th Jun 2024 AI Alignment Forum Linkpost for arxiv.org 14 min read 10 89 Ω 51 The paper is by Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov and David Krueger. This post was written by Fabien and Ryan, and may not reflect the views of Dmitrii and David. Scheming models might try to perform less capably than they are able to ( sandbag ). They migh
Explore this link on the map →related reading
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- AI in 2025: gestalt — LessWronglesswrong.com
- Memorizing weak examples can elicit strong behavior out of password-locked models — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hackingblog.redwoodresearch.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Model evals for dangerous capabilities — LessWronglesswrong.com
- Auditing Games for Sandbagging [paper] — LessWronglesswrong.com
- Fuzzing LLMs sometimes makes them reveal their secrets — AI Alignment Forumalignmentforum.org
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- 2312.06942arxiv.org