Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking
blog.redwoodresearch.org · 4,732 words · saved by 1 readers
A new analysis of the risk of AIs intentionally performing poorly.
Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking A new analysis of the risk of AIs intentionally performing poorly. Julian Stastny and Buck Shlegeris May 08, 2025 13 Share In the future, we will want to use powerful AIs on critical tasks such as doing AI safety and security R&D, dangerous capability evaluations, red-teaming safety protocols, or monitoring other powerful models. Since we care about models performing well on these tasks, we are worried about sandbagging: that if our models are misaligned 1 , they will intentionally underperform. San
related reading
- Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hackingredwoodresearch.substack.com
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Automated Researchers Can Subtly Sandbagalignment.anthropic.com
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- Reading Listblog.redwoodresearch.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Alignment Faking Mitigationsalignment.anthropic.com