2025.acl-long.453.pdf
aclanthology.org · 6,118 words · saved by 1 readers
N/A
Byte Latent Transformer: Patches Scale Better Than Tokens Artidoro Pagnoni1 , Ram Pasunuru‡ , Pedro Rodriguez‡ , John Nguyen‡ , Benjamin Muller, Margaret Li1 , Chunting Zhou⋄ , Lili Yu, Jason Weston, Luke Zettlemoyer3 , Gargi Ghosh, Mike Lewis, Ari Holtzman2,⋄,† , Srinivasan Iyer† FAIR at Meta, 1 Paul G. Allen School of Computer Science & Engineering, University of Washington, 2 University of Chicago…
saved by
related reading
- Byte Latent Transformer: Patches Scale Better Than Tokensarxiv.org
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- Byte Latent Transformer: Patches Scale Better Than Tokens | Research - AI at Metaai.meta.com
- The End of Tokenizationronaldyu.substack.com
- How To Scale Your Modeljax-ml.github.io
- MambaByte: Token-free Selective State Space Modelarxiv.org
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- The Annotated Transformernlp.seas.harvard.edu