flâneur — a map of the web's best reading

Anthropic's Pilot Sabotage Risk Report

alignment.anthropic.com · 973 words · saved by 1 readers

As practice for potential future Responsible Scaling Policy obligations, we're releasing a report on misalignment risk posed by our deployed models as of Summer 2025. We conclude that there is very low, but not fully negligible, risk of misaligned autonomous actions that substantially contribute to later catastrophic outcomes. We also release two reviews of this report: an internal review and an independent review by METR. Our Responsible Scaling Policy has so far come into force primarily to address risks related to high-stakes human misuse. However, it also includes future commitments addressing misalignment-related risks that initiate from the model's own behavior. For future models that pass a capability threshold we have not yet reached, it commits us to developing Affirmative cases of this kind are uncharted territory for the field: While we have published sketches of arguments we might use, we have never prepared a complete affirmative case of this kind, and are not aware of any

Anthropic's Pilot Sabotage Risk Report Alignment Science Blog Anthropic's Pilot Sabotage Risk Report Main Report: Samuel R. Bowman, Misha Wagner, Fabien Roger, and Holden Karnofsky Internal Review: Daniel M. Ziegler and Evan Hubinger October 28, 2025 tl;dr As practice for potential future Responsible Scaling Policy obligations, we're releasing a report on misalignment risk posed by our deployed models as of Summer 2025 . We conclude that there is very low, but not fully negligible, risk of misaligned autonomous actions that substantially contribute to later catastrophic outcomes. We also relea

Explore this link on the map →

related reading