How AI Is Learning to Think in Secret — LessWrong
lesswrong.com · 9,352 words · saved by 1 readers
On Thinkish, Neuralese, and the End of Readable Reasoning • ---------------------------------------- …
x How AI Is Learning to Think in Secret — LessWrong Language Models (LLMs) AI Curated 2026 Top Fifty: 21 % 386 How AI Is Learning to Think in Secret by Nicholas Andresen 6th Jan 2026 AI Alignment Forum Linkpost for nickandresen.substack.com 22 min read 33 386 Ω 56 On Thinkish, Neuralese, and the End of Readable Reasoning In September 2025, researchers released the internal monologue of OpenAI’s GPT-o3 as it decided to lie about scientific data. Here's what it was thinking: "We can glean disclaim disclaim synergy customizing illusions"? Pardon? This reads like someone had a stroke during a meet
related reading
- How AI Is Learning to Think in Secretnickandresen.substack.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- What’s your AI thinking? - AI Digesttheaidigest.org
- [2507.11473] Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safetyarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- As Rocks May Think | Eric Jangevjang.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org
- What is neuralese and why is everyone so concerned about it?transformernews.ai