Test-Time Training with KV Binding Is Secretly Linear Attention
research.nvidia.com · 1,091 words · saved by 1 readers
TTT with KV Binding Is Secretly Linear Attention
1 NVIDIA 2 University of Toronto 3 Vector Institute 4 Technion * Equal Contribution ICML 2026 TL;DR Test-time training (TTT) with KV binding is not test-time memorization — it's learned linear attention with enhanced representational capacity. (New to TTT?) Test-time training (TTT) with KV binding as sequence modeling layer is commonly interpreted as a form of online meta-learning that memorizes a key–value mapping at test time. However, our analysis reveals multiple phenomena that contradict this memorization-based interpretation. Motivated by these findings, we revisit the formulation…
saved by
related reading
- Learning to (Learn at Test Time): RNNs with Expressive Hidden Statesarxiv.org
- [2501.00663] Titans: Learning to Memorize at Test Timearxiv.org
- DeltaNet Explained (Part I) | Songlin Yangsustcsonglin.github.io
- Why test-time training? – Rabbitholessarahpannn.github.io
- [2512.23675] End-to-End Test-Time Training for Long Contextarxiv.org
- Linear Attention Fundamentals | Hailey Schoelkopfhaileyschoelkopf.github.io
- ali (@waterloo_intern) on Xx.com
- Test-Time Training Done Rightarxiv.org
- [2407.04620] Learning to (Learn at Test Time): RNNs with Expressive Hidden Statesar5iv.labs.arxiv.org
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- Log-Linear Attentionarxiv.org
- Kimi Linear: An Expressive, Efficient Attention Architecturearxiv.org