(Some of) The Models, They Just Don't Want to Learn | Tilde
Today's reasoning models usually scale sequential computation by generating more tokens. This works remarkably well, but it forces some intermediate computation to be carried out through generated tokens, rather than updated entirely within the model's latent state. This restriction can artificially serialize computations that could otherwise be performed in parallel within a latent representation. Alternative architectures could instead scale computation through additional latent depth as problems become harder. However, these models have often proven much harder to train. The difficulty may not be that better architectures for serial computation do not exist. It may be that our current knowledge of optimization disproportionately favors standard Transformers. We are launching One Layer Deeper, a competition for studying whether architectures, objectives and optimizers can be co-designed to learn deeper serial computation and productively extrapolate beyond the reasoning depths seen d
Back (Some of) The Models, They Just Don't Want to Learn 8.03.2026 × Correspondence to Sean McLeish, Ben Keigwin, Mark Saroufim, Rohan Anil Questions and discussion on Discord cite ↓ TL;DR Today's reasoning models usually scale sequential computation by generating more tokens. This works remarkably well, but it forces some intermediate computation to be carried out through generated tokens, rather than updated entirely within the model's latent state. This restriction can artificially serialize computations that could otherwise be performed in parallel within a latent representation. Alternati
saved by
related reading
- A Little Depth Goes a Long Way: The Expressive Power of Log-Depth Transformersarxiv.org
- The Little Book of Deep Learningfleuret.org
- Quantifying the Necessity of Chain of Thought through Opaque Serial Deptharxiv.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- As Rocks May Think | Eric Jangevjang.com
- Clouded Judgement 9.4.26 - Recurrent Depthcloudedjudgement.substack.com
- [2507.12549] The Serial Scaling Hypothesisarxiv.org
- An Alternative to Test-Time Scalingrentry.org
- Composer2.pdfcursor.com
- 2402.12875arxiv.org
- Memory makes computation universal, remember?thinks.lol
- the-illusion-of-thinking.pdfml-site.cdn-apple.com