[2512.19428] Attention Is Not What You Need
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status
[2512.19428] Attention Is Not What You Need --> Computer Science > Machine Learning arXiv:2512.19428 (cs) [Submitted on 22 Dec 2025] Title: Attention Is Not What You Need Authors: Zhang Chong View a PDF of the paper titled Attention Is Not What You Need, by Zhang Chong View PDF HTML (experimental) Abstract: We revisit a basic question in sequence modeling: is explicit self-attention actually necessary for strong performance and reasoning? We argue that standard multi-head attention is best seen as a form of tensor lifting: hidden vectors are mapped into a high-dimensional space of pairwise int
Explore this link on the map →saved by
related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- transformer_attention.pdfarxiv.org
- 1706.03762arxiv.org
- [1706.03762] Attention Is All You Needarxiv.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Transformers from Scratche2eml.school
- Overleaf Examplearxiv.org
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Everything About Transformerskrupadave.com
- Mamba: The Easy Wayjackcook.com