(Some of) The Models, They Just Don't Want to Learn | Tilde
Today's reasoning models usually scale sequential computation by generating more tokens. This works remarkably well, but it forces some intermediate computation to be carried out through generated tokens, rather than updated entirely within the model's latent state. This restriction can artificially serialize computations that could otherwise be performed in parallel within a latent representation. Alternative architectures could instead scale computation through additional latent depth as problems become harder. However, these models have often proven much harder to train. The difficulty may not be that better architectures for serial computation do not exist. It may be that our current knowledge of optimization disproportionately favors standard Transformers. We are launching One Layer Deeper, a competition for studying whether architectures, objectives and optimizers can be co-designed to learn deeper serial computation and productively extrapolate beyond the reasoning depths seen d
Back (Some of) The Models, They Just Don't Want to Learn 8.03.2026 × Correspondence to Sean McLeish, Ben Keigwin, Mark Saroufim, Rohan Anil Questions and discussion on Discord cite ↓ TL;DR Today's reasoning models usually scale sequential computation by generating more tokens. This works remarkably well, but it forces some intermediate computation to be carried out through generated tokens, rather than updated entirely within the model's latent state. This restriction can artificially serialize computations that could otherwise be performed in parallel within a latent representation. Alternati
Explore this link on the map →saved by
related reading
- The Decade of Deep Learning | Leo Gaobmk.sh
- The Little Book of Deep Learningfleuret.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- How To Scale Your Modeljax-ml.github.io
- An Alternative to Test-Time Scalingrentry.org
- Composer2.pdfcursor.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Memory makes computation universal, remember?thinks.lol
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- [2402.12875] Chain of Thought Empowers Transformers to Solve Inherently Serial Problemsarxiv.org
- [2602.05970] Inverse Depth Scaling From Most Layers Being Similararxiv.org