Byte Latent Transformer: Patches Scale Better Than Tokens
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
]FAIR at Meta 1]Paul G. Allen School of Computer Science & Engineering, University of Washington 2]University of Chicago \contribution [‡]Joint second author \contribution [†]Joint last author \contribution [⋄]Work done at Meta Byte Latent Transformer: Patches Scale Better Than Tokens Artidoro Pagnoni Ram Pasunuru Pedro Rodriguez John Nguyen Benjamin Muller Margaret Li Chunting Zhou Lili Yu Jason Weston Luke Zettlemoyer Gargi Ghosh Mike Lewis Ari Holtzman Srinivasan Iyer [ [ [ cs.washington.edu meta.com (July 25, 2025) Abstract We introduce the Byte Latent Transformer ( B LT), a new byte-level
Explore this link on the map →saved by
related reading
- MambaByte: Token-free Selective State Space Modelarxiv.org
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- Byte Latent Transformer: Patches Scale Better Than Tokens | Research - AI at Metaai.meta.com
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- The Annotated Transformernlp.seas.harvard.edu
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- The Annotated Transformernlp.seas.harvard.edu
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Transformers from Scratche2eml.school
- 1706.03762arxiv.org