flâneur — a map of the web's best reading

The Mamba Effect: State Space Models Taking on Transformers

hungleai.substack.com · 4,748 words · saved by 1 readers

Large Language Models (LLMs) are pretrained on massive datasets to achieve AGI (Artificial General Intelligence). As an unwritten rule, the Transformer [9] architecture is the backbone of LLMs due to its ability to capture rich representations through attention layers. These layers provide direct access to past inputs at any point during processing. However, this capability comes with a computational cost of O(L2) complexity, where L is the number of timesteps (tokens) the Transformer needs to process. With the development of advanced GPUs and significant investment in AI infrastructure, the quadratic complexity of Transformers has become less of a barrier, allowing Transformer-based LLMs (like OpenAI's ChatGPT, Anthropic's Claude, and Google's Gemini) to achieve excellent results and gain widespread popularity in many real-world applications. However, there are concerns about the Transformer approach: ❌ Transformer intelligence, driven by attention mechanisms, is artificial and doesn'

Spotlight The Mamba Effect: State Space Models Taking on Transformers Mamba: Linear-Time Sequence Modeling with Selective State Spaces (Outstanding Paper Award, COLM 2024) Hung Le Jul 03, 2024 4 1 Share Table of Content Large Language Models, Transformers, and the Fundamental Bottleneck Mamba Dissection: A Top-Down Approach Linear-Time Decoding State Space Model Foundation Selective State Spaces Mamba Empirical Performance Mamba is Faster than Transformers Mamba Scales Linearly up to a Million Tokens State-of-the-art Performance Rivaling Transformers in Language Modeling Ablation Studies Final

Explore this link on the map →

related reading