What is neuralese and why is everyone so concerned about it?
transformernews.ai · 1,345 words · saved by 1 readers
OpenAI’s new Astra model is raising concerns about our continued ability to monitor AI’s chain of thought
Image: Jacob Wackerhausen / Getty Images Off the back of a report in The Information on Tuesday, everyone’s suddenly very worried about “neuralese”. OpenAI’s new Astra model, the outlet reported, was built with a new architecture which could make it harder to monitor its reasoning, raising safety concerns. Ryan Greenblatt, a prominent AI safety researcher, said that if true, the news “may be the single worst development for AI security/safety to date”. But what does any of this mean — and is the latest development really as bad as some have made it out to be? Let’s start with the very…
saved by
related reading
- Astra Is Hard to Monitorthezvi.substack.com
- AI #184: Post Post Mortemthezvi.substack.com
- How AI Is Learning to Think in Secretnickandresen.substack.com
- AI 2027ai-2027.com
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- Neel Nanda on the race to read AI minds (part 1) | 80,000 Hours80000hours.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- Policy Options for Preserving Chain of Thought Monitorability — Institute for AI Policy and Strategyiaps.ai
- AI 2027ai-2027.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com