✳flâneur — a map of the web's best reading
Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking
blog.redwoodresearch.org · 4,732 words · saved by 1 readers
A new analysis of the risk of AIs intentionally performing poorly.
Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking A new analysis of the risk of AIs intentionally performing poorly. Julian Stastny and Buck Shlegeris May 08, 2025 13 Share In the future, we will want to use powerful AIs on critical tasks such as doing AI safety and security R&D, dangerous capability evaluations, red-teaming safety protocols, or monitoring other powerful models. Since we care about models performing well on these tasks, we are worried about sandbagging: that if our models are misaligned 1 , they will intentionally underperform. San
Explore this link on the map →related reading
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Automated Researchers Can Subtly Sandbagalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com