Astra Is Hard to Monitor - by Zvi Mowshowitz
thezvi.substack.com · 10,008 words · saved by 1 readers
OpenAI’s central message on Astra is that it is three things:
OpenAI’s central message on Astra is that it is three things: Highly capable and can do all the things for you. Hard to monitor. The most aligned model. The first claim largely checks out. Astra and Fable are both clearly excellent models. This post is about their second claim, which to their credit they are being loud about, in three parts: The system card result, affirmed on Twitter by several OpenAI employees including Tomek Korbak, and in an excellent post by Chief Scientist Jakub Pachocki that I covered yesterday, that Astra is harder to monitor. OpenAI’s use of recurrent depth…
saved by
related reading
- What is neuralese and why is everyone so concerned about it?transformernews.ai
- AI #184: Post Post Mortemthezvi.substack.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Proposal for tracking the effects of architecture on monitorability — Redwood Researchredwoodresearch.org
- AI 2027ai-2027.com
- [2512.18311] Monitoring Monitorabilityarxiv.org
- [2512.18311] Monitoring Monitorabilityarxiv.org
- Policy Options for Preserving Chain of Thought Monitorability — Institute for AI Policy and Strategyiaps.ai
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org