Early work on monitorability evaluations - METR
We show preliminary results on a prototype evaluation that tests monitors' ability to catch AI agents doing side tasks, and AI agents' ability to bypass this monitoring.
Early work on monitorability evaluations - METR Our Work Research Notes Updates Risk Assessment About Donate Careers Search --> Our Work Research Notes Updates Risk Assessment About Donate Careers Menu × Early work on monitorability evaluations CONTRIBUTORS Megan Kinniment , Seraphina Nix , Thomas Broadley , Hjalmar Wijk , and Neev Parikh DATE January 22, 2026 SHARE Copy Link Citation BibTeX Citation × @misc { metr-2026-early-work-on-monitorability-evaluations , title = {Early work on monitorability evaluations} , author = {Megan Kinniment, Seraphina Nix, Thomas Broadley, Hjalmar W
Explore this link on the map →saved by
related reading
- MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity - METRmetr.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Demystifying evals for AI agents \ Anthropicanthropic.com
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- Measuring AI Ability to Complete Long Tasks - METRmetr.org
- Open Sourcing Monitorability Evaluationsalignment.openai.com
- Agentic Monitoring for AI Control — LessWronglesswrong.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Agent Observability and Tracingarize.com
- 2312.06942arxiv.org