Logit ROCs: Monitor TPR is linear in FPR in logit space
substack.com · 3,720 words · saved by 1 readers
Summary
We study trusted monitoring for AI control, where a weaker trusted model reviews the actions of a stronger untrusted agent and flags suspicious behavior for human audit. We propose a simple mathematical model relating safety (true positive rate) to audit budget (false positive rate) at realistic, low FPRs: logit(TPR) is linear in logit(FPR). Equivalently, benign and attack scores can be modeled with logistic distributions. This gives two practical benefits. It lets practitioners estimate TPRs more accurately at deployment-relevant FPRs when data are limited, and it gives a compact,…
related reading
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- 2312.06942arxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- Why We Are Excited About Confessionsalignment.openai.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Open Sourcing Monitorability Evaluationsalignment.openai.com
- Agentic Monitoring for AI Control — LessWronglesswrong.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Early work on monitorability evaluations - METRmetr.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org