flâneur

Test-Time Training with KV Binding Is Secretly Linear Attention

research.nvidia.com · 1,091 words · saved by 1 readers

TTT with KV Binding Is Secretly Linear Attention

1 NVIDIA 2 University of Toronto 3 Vector Institute 4 Technion * Equal Contribution ICML 2026 TL;DR Test-time training (TTT) with KV binding is not test-time memorization — it's learned linear attention with enhanced representational capacity. (New to TTT?) Test-time training (TTT) with KV binding as sequence modeling layer is commonly interpreted as a form of online meta-learning that memorizes a key–value mapping at test time. However, our analysis reveals multiple phenomena that contradict this memorization-based interpretation. Motivated by these findings, we revisit the formulation…

saved by

related reading