From Pairwise to Higher Order Tensor Operations on GPUs – Springtail Blog
Pairwise primitives are a primary operational pattern in deep learning. They take two inputs and fuse them into a single output. Matrix multiplication, dot product, and element-wise or Hadamard product are all examples of this. In transformers, self-attention computes relationships between pairs of tokens. Similarly, gating mechanisms such as Gated Linear Units (GLUs) combine two signals via component-wise multiplication – one vector gates or modulates another. These are extremely useful for capturing complex relationships in high dimensional space. For example, self-attention can capture long-range dependencies in input sequences, and gating units stabilize and enhance representational power of neural network learning. However, pairwise primitives are by definition limited to two inputs, treating interactions as linear or bilinear combinations. To capture higher-order dependencies between three or more inputs, networks must compose sequences of pairwise operations to approximate n-way
From Pairwise to Higher Order Tensor Operations on GPUs – Springtail Blog From Pairwise to Higher Order Tensor Operations on GPUs From Pairwise to Higher Order Tensor Operations on GPUs pdf version of this post Anosha Rahim & Timothy Hanson September 30, 2025 Pairwise primitives are a primary operational pattern in deep learning. They take two inputs and fuse them into a single output. Matrix multiplication, dot product, and element-wise or Hadamard product are all examples of this. In transformers, self-attention computes relationships between pairs of tokens. Similarly, gating mechanis
Explore this link on the map →saved by
related reading
- Inside NVIDIA GPUs: Anatomy of high performance matmul kernels - Aleksa Gordićaleksagordic.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- CVPR2023_eff_tutorial_molchanov.pdfnvlabs.github.io
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Overleaf Examplearxiv.org
- Making Deep Learning go Brrrr From First Principleshorace.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- From Online Softmax to FlashAttentioncourses.cs.washington.edu
- GPUs Go Brrr · Hazy Researchhazyresearch.stanford.edu
- GPU Performance Background User's Guide - NVIDIA Docsdocs.nvidia.com
- The Best GPUs for Deep Learning in 2023 — An In-depth Analysistimdettmers.com