Proposal for tracking the effects of architecture on monitorability — Redwood Research
Architectures that incorporate opaque recurrence or allow agents to communicate using latents could rapidly make it much harder to monitor chains of thought. We propose that AI companies regularly report verified information about opaque serial depth, share monitorability evidence, and publish a policy on architectures that could degrade monitorability.
Architectures that incorporate opaque recurrence or allow for agents to communicate with each other using latents could rapidly make it much harder to monitor chains of thought or communication (we’ll refer to this property as “monitorability” going forward).1 As companies begin to explore such architectures, we believe it is important to transparently share evidence about how monitorability varies with architecture and training method. To inform the scientific debate on how to make tradeoffs between performance and monitorability, we believe AI companies should: Regularly report externally…
saved by
related reading
- Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settingsarxiv.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- Astra Is Hard to Monitorthezvi.substack.com
- [2512.18311] Monitoring Monitorabilityarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- Early work on monitorability evaluations - METRmetr.org
- Policy Options for Preserving Chain of Thought Monitorability — Institute for AI Policy and Strategyiaps.ai
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- Quantifying the Necessity of Chain of Thought through Opaque Serial Deptharxiv.org