Towards a Typology of Strange LLM Chains-of-Thought
LLMs being trained with RLVR (Reinforcement Learning from Verifiable Rewards) start off with a 'chain-of-thought' (CoT) in whatever language the LLM was originally trained on. But after a long period of training, the CoT sometimes starts to look very weird; to resemble no human language; or even to grow completely unintelligible. Why might this happen? I've seen a lot of speculation about why. But a lot of this speculation narrows too quickly, to just one or two hypotheses. My intent is also to speculate, but more broadly. Specifically, I want to outline six nonexclusive possible causes for the weird tokens: new better language, spandrels, context refresh, deliberate obfuscation, natural drift, and conflicting shards. And I also wish to outline ideas for experiments and evidence that could help us distinguish these causes. I'm sure I'm not enumerating the full space of possibilities. I'm also sure that I'm probably making some mistakes what follows, or confusing my ontologies. But it's
Towards a Typology of Strange LLM Chains-of-Thought 1 A 3 O R N 1 A 3 O R N Towards a Typology of Strange LLM Chains-of-Thought On the Causes of LLM Unintelligibility Created: 2025-10-02 Wordcount: 2k Tags: philosophy effort-post machine-learning empiricism Intro LLMs being trained with RLVR (Reinforcement Learning from Verifiable Rewards) start off with a 'chain-of-thought' (CoT) in whatever language the LLM was originally trained on. But after a long period of training, the CoT sometimes starts to look very weird; to resemble no human language; or even to grow completely unintelligible. Why
saved by
related reading
- The Unintelligibility is Ours: Notes on Chain of Thought1a3orn.com
- Towards a Typology of Strange LLM Chains-of-Thought — LessWronglesswrong.com
- Training Large Language Models to Reason in a Continuous Latent Spacearxiv.org
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Simulated Users & Sad LLMs1a3orn.com
- Prompt Injection as Role Confusionrole-confusion.github.io
- LLM Daydreaming · Gwern.netgwern.net
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Chain of Draft: Thinking Faster by Writing Lessarxiv.org
- dennyzhou.github.io/LLM-Reasoning-Stanford-CS-25.pdfdennyzhou.github.io
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org