Using group theory to explore the space of positional encodings for attention
Attention is a computational primitive at the core of modern language models, allowing internal representations to reference and influence each other. It’s h...
Jane Street Blog - Using group theory to explore the space of positional encodings for attention Using group theory to explore the space of positional encodings for attention Apr 22, 2026 | 13 min read Share on Facebook Share on Twitter Share on LinkedIn By: Alok Puranik Attention is a computational primitive at the core of modern language models, allowing internal representations to reference and influence each other. It’s how these models handle sequential data in the first place. Yet, naively implemented, attention doesn’t have any notion of position. In the core attention computation, you
Explore this link on the map →saved by
related reading
- Your Transformer is Secretly an EOT Solver | Elements of a Vector Spaceelonlit.com
- Master Positional Encoding: Part I | Towards Data Sciencetowardsdatascience.com
- On N-dimensional Rotary Positional Embeddingsjerryxio.ng
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Relative Positional Encoding - Jake Taejaketae.github.io
- Transformer Architecture: The Positional Encoding - Amirhossein Kazemnejad's Blogkazemnejad.com
- Rotary Embeddings: A Relative Revolution | EleutherAI Blogblog.eleuther.ai
- Transformers from Scratche2eml.school
- transformer_attention.pdfarxiv.org
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io