Pengyu Zhao on X: "MiniMax M2 Tech Blog 3: Why Did M2 End Up as a Full Attention Model? On behave of pre-training lead Haohai Sun. (https://t.co/WH4xOD9KrT) I. Introduction As the lead of MiniMax-M2 pretrain, I've been getting many queries from the community on "Why did you turn back the clock" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Messages Home Explore 1 Notifications Messages SuperGrok Premium Lists Bookmarks Communities Profile More Post aaron @aarnphm_ Post Reply See new posts Conversation Pengyu Zhao @zpysky1125 MiniMax M2 Tech Blog 3: Why Did M2 End Up as a Full Attention Model? On behave of pre-training lead Haohai Sun. ( https:// zhihu.com/question/19653 02088260104295/answer/1966810157473335067 …) I. Introduction As the lead of MiniMax-M2 pretrain, I've been getting many queries from the community on "Why did you turn back the clock and go with full attention with MiniMax M2?" After explaining the backstory in one chat after another, I figured it's time to write down our journey in a blog. Honestly, I could give you the textbook debate. I could talk all afternoon about why you should build linear/sparse attention. Then, I could turn around and talk all afternoon about why you shouldn't. But what's the point of all that hand-waving?
Pengyu Zhao @zpysky1125 MiniMax M2 Tech Blog 3: Why Did M2 End Up as a Full Attention Model? On behave of pre-training lead Haohai Sun. ( zhihu.com/question/19653… ) I. Introduction As the lead of MiniMax-M2 pretrain, I've been getting many queries from the community on "Why did you turn back the clock and go with full attention with MiniMax M2?" After explaining the backstory in one chat after another, I figured it's time to write down our journey in a blog. Honestly, I could give you the textbook debate. I could talk all afternoon about why you should build linear/sparse attention. Then, I c
Explore this link on the map →related reading
- 2502.11089arxiv.org
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Mamba: The Easy Wayjackcook.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Linear Attention Is All You Need | Towards Data Sciencetowardsdatascience.com
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Composer2.pdfcursor.com
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev