MambaByte: Token-free Selective State Space Model
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
\newfloatcommand capbtabboxtable[][ \FBwidth ] MambaByte: Token-free Selective State Space Model Junxiong Wang, Tushaar Gangavarapu, Jing Nathan Yan, Alexander M. Rush Cornell University { jw2544 , tg352 , jy858 , arush }@cornell.edu Abstract Token-free language models learn directly from raw bytes and remove the inductive bias of subword tokenization. Operating on bytes, however, results in significantly longer sequences. In this setting, standard autoregressive Transformers scale poorly as the effective memory required grows with sequence length. The recent Mamba state space model (SSM) deve
Explore this link on the map →saved by
related reading
- Byte Latent Transformer: Patches Scale Better Than Tokensarxiv.org
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- A Visual Guide to Mamba and State Space Modelsnewsletter.maartengrootendorst.com
- Mamba: The Easy Wayjackcook.com
- Mamba Explainedthegradient.pub
- Mamba No. 5 (A Little Bit Of…) | Sparse Notesjameschen.io
- H-Nets - the Past | Goomba Labgoombalab.github.io
- H3: Language Modeling with State Space Models and (Almost) No Attention · Hazy Researchhazyresearch.stanford.edu
- The Mamba Effect: State Space Models Taking on Transformershungleai.substack.com
- A Visual Guide to Mamba and State Space Models - Maarten Grootendorstmaartengrootendorst.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io