The Mamba Effect: State Space Models Taking on Transformers
Large Language Models (LLMs) are pretrained on massive datasets to achieve AGI (Artificial General Intelligence). As an unwritten rule, the Transformer [9] architecture is the backbone of LLMs due to its ability to capture rich representations through attention layers. These layers provide direct access to past inputs at any point during processing. However, this capability comes with a computational cost of O(L2) complexity, where L is the number of timesteps (tokens) the Transformer needs to process. With the development of advanced GPUs and significant investment in AI infrastructure, the quadratic complexity of Transformers has become less of a barrier, allowing Transformer-based LLMs (like OpenAI's ChatGPT, Anthropic's Claude, and Google's Gemini) to achieve excellent results and gain widespread popularity in many real-world applications. However, there are concerns about the Transformer approach: ❌ Transformer intelligence, driven by attention mechanisms, is artificial and doesn'
Spotlight The Mamba Effect: State Space Models Taking on Transformers Mamba: Linear-Time Sequence Modeling with Selective State Spaces (Outstanding Paper Award, COLM 2024) Hung Le Jul 03, 2024 4 1 Share Table of Content Large Language Models, Transformers, and the Fundamental Bottleneck Mamba Dissection: A Top-Down Approach Linear-Time Decoding State Space Model Foundation Selective State Spaces Mamba Empirical Performance Mamba is Faster than Transformers Mamba Scales Linearly up to a Million Tokens State-of-the-art Performance Rivaling Transformers in Language Modeling Ablation Studies Final
Explore this link on the map →related reading
- A Visual Guide to Mamba and State Space Modelsnewsletter.maartengrootendorst.com
- Mamba Explainedthegradient.pub
- A Visual Guide to Mamba and State Space Models - Maarten Grootendorstmaartengrootendorst.com
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Mamba: The Easy Wayjackcook.com
- Mamba No. 5 (A Little Bit Of…) | Sparse Notesjameschen.io
- MambaByte: Token-free Selective State Space Modelarxiv.org
- H3: Language Modeling with State Space Models and (Almost) No Attention · Hazy Researchhazyresearch.stanford.edu
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Transformers from Scratche2eml.school
- State Space Duality (Mamba-2) Part I - The Model | Goomba Labgoombalab.github.io