Open Sourcing Monitorability Evaluations
We open-source datasets and code from our Monitoring Monitorability paper, and share a new filtering strategy for noise-dominated intervention evaluation instances.
Open Sourcing Monitorability Evaluations ← Back to OpenAI Alignment Blog Open Sourcing Monitorability Evaluations Apr 23, 2026 · Melody Y. Guan, Miles Wang, Micah Carroll, Zehao Dou, Annie Y. Wei, Marcus Williams, Benjamin Arnav, Joost Huizinga, Ian Kivlichan, Mia Glaese, Jakub Pachocki, Bowen Baker TL;DR We are releasing a subset of datasets and reference code from our chain-of-thought monitorability work . The release includes most datasets from our monitorability evaluation suite, code for computing the monitorability metric g-mean 2 , and a new cross-fit filtering strategy that makes
Explore this link on the map →saved by
related reading
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- gpt-4.pdfcdn.openai.com
- [2512.18311] Monitoring Monitorabilityarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Early work on monitorability evaluations - METRmetr.org
- 2312.06942arxiv.org
- Demystifying evals for AI agents \ Anthropicanthropic.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- [2603.05706] Reasoning Models Struggle to Control their Chains of Thoughtarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org