H3: Language Modeling with State Space Models and (Almost) No Attention · Hazy Research
hazyresearch.stanford.edu · 1,279 words · saved by 5 readers
Replacing attention with SSMs in language modeling.
H3: Language Modeling with State Space Models and (Almost) No Attention · Hazy Research Jan 23, 2023 · 7 min read H3: Language Modeling with State Space Models and (Almost) No Attention Dan Fu , Tri Dao , Khaled Saab , Armin Thomas , Atri Rudra , and Chris Ré . State space models (SSMs) are strong general-purpose sequence models, but have underperformed attention in language modeling. How can we close this gap? In this blog post, we’ll take a look at some critical capability gaps – using some simple synthetic languages of in-context as a guide. We’ll use our understanding to build H3 (Hungry H
saved by
related reading
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- [2212.14052] Hungry Hungry Hippos: Towards Language Modeling with State Space Modelsarxiv.org
- Mamba Explainedthegradient.pub
- A Visual Guide to Mamba and State Space Modelsnewsletter.maartengrootendorst.com
- 2312.00752arxiv.org
- Mamba: The Easy Wayjackcook.com
- [2111.00396] Efficiently Modeling Long Sequences with Structured State Spacesarxiv.org
- 2405.21060-Transformers are SSMs (Mamba2)arxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- A History of Large Language Modelsgregorygundersen.com
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Log-Linear Attentionarxiv.org