✳flâneur — a map of the web's best reading
H3: Language Modeling with State Space Models and (Almost) No Attention · Hazy Research
hazyresearch.stanford.edu · 1,279 words · saved by 5 readers
Replacing attention with SSMs in language modeling.
H3: Language Modeling with State Space Models and (Almost) No Attention · Hazy Research Jan 23, 2023 · 7 min read H3: Language Modeling with State Space Models and (Almost) No Attention Dan Fu , Tri Dao , Khaled Saab , Armin Thomas , Atri Rudra , and Chris Ré . State space models (SSMs) are strong general-purpose sequence models, but have underperformed attention in language modeling. How can we close this gap? In this blog post, we’ll take a look at some critical capability gaps – using some simple synthetic languages of in-context as a guide. We’ll use our understanding to build H3 (Hungry H
Explore this link on the map →saved by
related reading
- [2212.14052] Hungry Hungry Hippos: Towards Language Modeling with State Space Modelsarxiv.org
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Mamba Explainedthegradient.pub
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- A Visual Guide to Mamba and State Space Modelsnewsletter.maartengrootendorst.com
- Mamba: The Easy Wayjackcook.com
- Some Intuition on Attention and the Transformereugeneyan.com
- A Visual Guide to Mamba and State Space Models - Maarten Grootendorstmaartengrootendorst.com
- transformer_attention.pdfarxiv.org
- The Mamba Effect: State Space Models Taking on Transformershungleai.substack.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- [2111.00396] Efficiently Modeling Long Sequences with Structured State Spacesarxiv.org