flâneur — a map of the web's best reading

Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking

blog.redwoodresearch.org · 4,732 words · saved by 1 readers

A new analysis of the risk of AIs intentionally performing poorly.

Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking A new analysis of the risk of AIs intentionally performing poorly. Julian Stastny and Buck Shlegeris May 08, 2025 13 Share In the future, we will want to use powerful AIs on critical tasks such as doing AI safety and security R&D, dangerous capability evaluations, red-teaming safety protocols, or monitoring other powerful models. Since we care about models performing well on these tasks, we are worried about sandbagging: that if our models are misaligned 1 , they will intentionally underperform. San

Explore this link on the map →

related reading