flâneur — a map of the web's best reading

(Some of) The Models, They Just Don't Want to Learn | Tilde

blog.tilderesearch.com · 2,395 words · saved by 1 readers

Today's reasoning models usually scale sequential computation by generating more tokens. This works remarkably well, but it forces some intermediate computation to be carried out through generated tokens, rather than updated entirely within the model's latent state. This restriction can artificially serialize computations that could otherwise be performed in parallel within a latent representation. Alternative architectures could instead scale computation through additional latent depth as problems become harder. However, these models have often proven much harder to train. The difficulty may not be that better architectures for serial computation do not exist. It may be that our current knowledge of optimization disproportionately favors standard Transformers. We are launching One Layer Deeper, a competition for studying whether architectures, objectives and optimizers can be co-designed to learn deeper serial computation and productively extrapolate beyond the reasoning depths seen d

Back (Some of) The Models, They Just Don't Want to Learn 8.03.2026 × Correspondence to Sean McLeish, Ben Keigwin, Mark Saroufim, Rohan Anil Questions and discussion on Discord cite ↓ TL;DR Today's reasoning models usually scale sequential computation by generating more tokens. This works remarkably well, but it forces some intermediate computation to be carried out through generated tokens, rather than updated entirely within the model's latent state. This restriction can artificially serialize computations that could otherwise be performed in parallel within a latent representation. Alternati

Explore this link on the map →

saved by

related reading