The Bitter Lesson is coming for Tokenization | ⛰️ lucalp
Highlights the desire to replace tokenization with a general method that better leverages compute and data. We'll see tokenization's fragility and review the Byte Latent Transformer arch.
The Bitter Lesson is coming for Tokenization – ⛰️ lucalp The Bitter Lesson is coming for Tokenization 24 Jun, 2025 a world of LLMs without tokenization is desirable and increasingly possible Published on 24/06/2025 • ⏱️ 29 min read In this post, we highlight the desire to replace tokenization with a general method that better leverages compute and data. We'll see tokenization's role, its fragility and we'll build a case for removing it. After understanding the design space, we'll explore the potential impacts of a recent promising candidate (Byte Latent Transformer) and build strong intuitions
saved by
related reading
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- The End of Tokenizationronaldyu.substack.com
- Byte Latent Transformer: Patches Scale Better Than Tokensarxiv.org
- 2025.acl-long.453.pdfaclanthology.org
- [2112.10508] Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLParxiv.org
- Byte Latent Transformer: Patches Scale Better Than Tokens | Research - AI at Metaai.meta.com
- MambaByte: Token-free Selective State Space Modelarxiv.org
- 2502.11089arxiv.org
- H-Nets - the Past | Goomba Labgoombalab.github.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Annotated Transformernlp.seas.harvard.edu
- You Can Learn Tokenization End-to-End with Reinforcement Learningarxiv.org