How AI Is Learning to Think in Secret
nickandresen.substack.com · 5,447 words · saved by 3 readers
On Thinkish, Neuralese, and the End of Readable Reasoning
In September 2025, researchers released the internal monologue of OpenAI’s GPT-o3 as it decided to lie about scientific data. Here’s what it was thinking: “We can glean disclaim disclaim synergy customizing illusions”? Pardon? This reads like someone had a stroke during a meeting they didn’t want to be in, but their hand kept taking notes. The transcript comes from a recent paper by Apollo Research and OpenAI on catching AI systems scheming. To understand what’s happening here - and why one of the most sophisticated AI systems in the world is babbling about “disclaim disclaim vantage” - it…
saved by
related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- AI 2027ai-2027.com
- As Rocks May Think | Eric Jangevjang.com
- Learning to reason with LLMs | OpenAIopenai.com
- A global workspace in language models \ Anthropicanthropic.com
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- What is neuralese and why is everyone so concerned about it?transformernews.ai