Open Sourcing Monitorability Evaluations
We open-source datasets and code from our Monitoring Monitorability paper, and share a new filtering strategy for noise-dominated intervention evaluation instances.
Open Sourcing Monitorability Evaluations ← Back to OpenAI Alignment Blog Open Sourcing Monitorability Evaluations Apr 23, 2026 · Melody Y. Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y. Wei, Marcus Williams, Benjamin Arnav, Joost Huizinga, Ian Kivlichan, Mia Glaese, Jakub Pachocki, Bowen Baker TL;DR We are releasing a subset of datasets and reference code from our chain-of-thought monitorability work . The release includes most datasets from our monitorability evaluation suite, code for computing the monitorability metric g-mean 2 , and a new cross-fit filtering strategy that makes
saved by
related reading
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- 2312.06942arxiv.org
- Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settingsarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Early work on monitorability evaluations - METRmetr.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Proposal for tracking the effects of architecture on monitorability — Redwood Researchredwoodresearch.org
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org