Early work on monitorability evaluations - METR
We show preliminary results on a prototype evaluation that tests monitors' ability to catch AI agents doing side tasks, and AI agents' ability to bypass this monitoring.
Early work on monitorability evaluations - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Early work on monitorability evaluations CONTRIBUTORS Megan Kinniment , Seraphina Nix , Thomas Broadley , Hjalmar Wijk , and Neev Parikh DATE January 22, 2026 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2026-early-work-on-monitorability-evaluations , title = {Early work on monitorability evaluations} , author = {Megan Kinniment, Seraphina Nix, Thomas Broadley, Hjalmar W
saved by
related reading
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Demystifying evals for AI agents \ Anthropicanthropic.com
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- How to build fast, efficient monitors for AI models using probes - Goodfiregoodfire.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Proposal for tracking the effects of architecture on monitorability — Redwood Researchredwoodresearch.org
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Agentic Monitoring for AI Control — LessWronglesswrong.com