✳flâneur — a map of the web's best reading
On the Tradeoffs of SSMs and Transformers | Goomba Lab
goombalab.github.io · 6,672 words · saved by 8 readers
(or - tokens are bs)
On the Tradeoffs of SSMs and Transformers | Goomba Lab On the Tradeoffs of SSMs and Transformers (or - tokens are bs) This blog post was adapted from a talk I’ve given a handful of times over the last year. It was meant to be a high-level talk accessible to a fairly broad audience, but hopefully has some interesting insights, opinions, and intuitions around sequence models for the dedicated researchers too. State Space Models Just so we’re on the same page, I’ll start by defining what I mean by a state space model. (This section isn’t strictly necessary to get to the main part of this post tho
Explore this link on the map →saved by
related reading
- A Visual Guide to Mamba and State Space Modelsnewsletter.maartengrootendorst.com
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- H3: Language Modeling with State Space Models and (Almost) No Attention · Hazy Researchhazyresearch.stanford.edu
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Mamba Explainedthegradient.pub
- H-Nets - the Past | Goomba Labgoombalab.github.io
- The Mamba Effect: State Space Models Taking on Transformershungleai.substack.com
- A Visual Guide to Mamba and State Space Models - Maarten Grootendorstmaartengrootendorst.com
- transformer_attention.pdfarxiv.org
- Transformers from Scratche2eml.school
- [2212.14052] Hungry Hungry Hippos: Towards Language Modeling with State Space Modelsarxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io