CoT monitorability: why g-means and not F1?
substack.com · 1,142 words · saved by 1 readers
It all has to do with separating monitorability from misalignment property.
Figure 3 in “Monitoring Monitorability” (see below) argues why g-mean is better than F1. I originally didn’t get the implication (esp. about Case B figure) until I understood the mathematical differences between them, and what the implications are when it comes to monitoring misbehaving models. You have a binary monitor that says “misbehavior detected” or “no misbehavior.” The ground truth is either positive (bad behavior happened) or negative (it didn’t). The four outcomes are TP, FP, TN, FN as usual and there are 𝒩 samples, then we get the following operations: TPR = TP/(TP+FN) — “of…
saved by
related reading
- Why We Are Excited About Confessionsalignment.openai.com
- SOTA alignment assessments don’t strongly update us against misalignmentblog.redwoodresearch.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Open Sourcing Monitorability Evaluationsalignment.openai.com
- Precision and recallen.wikipedia.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Logit ROCs: Monitor TPR is linear in FPR in logit spacesubstack.com
- Early work on monitorability evaluations - METRmetr.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- A Mike's-Eye View of ARC's Research — Alignment Research Centeralignment.org
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org