[2510.26745] Deep sequence models tend to memorize geometrically; it is unclear why
Abstract:Deep sequence models are said to store atomic facts predominantly in the form of associative memory: a brute-force lookup of co-occurring entities. We identify a dramatically different form of storage of atomic facts that we term as geometric memory. Here, the model has synthesized embeddings encoding novel global relationships between all entities, including ones that do not co-occur in training. Such storage is powerful: for instance, we show how it transforms a hard reasoning task involving an $\ell$-fold composition into an easy-to-learn $1$-step navigation task. From this phenomenon, we extract fundamental aspects of neural embedding geometries that are hard to explain. We argue that the rise of such a geometry, as against a lookup of local associations, cannot be straightforwardly attributed to typical supervisory, architectural, or optimizational pressures. Counterintuitively, a geometry is learned even when it is more complex than the brute-force lookup. Then, by analyzing a connection to Node2Vec, we demonstrate how the geometry stems from a spectral bias that -- in contrast to prevailing theories -- indeed arises naturally despite the lack of various pressures. This analysis also points out to practitioners a visible headroom to make Transformer memory more strongly geometric. We hope the geometric view of parametric memory encourages revisiting the default intuitions that guide researchers in areas like knowledge acquisition, capacity, discovery, and unlearning.
[2510.26745] Deep sequence models tend to memorize geometrically; it is unclear why --> Computer Science > Machine Learning arXiv:2510.26745 (cs) [Submitted on 30 Oct 2025 ( v1 ), last revised 18 May 2026 (this version, v3)] Title: Deep sequence models tend to memorize geometrically; it is unclear why Authors: Shahriar Noroozizadeh , Vaishnavh Nagarajan , Elan Rosenfeld , Sanjiv Kumar View a PDF of the paper titled Deep sequence models tend to memorize geometrically; it is unclear why, by Shahriar Noroozizadeh and 3 other authors View PDF HTML (experimental) Abstract: Deep sequence models are
Explore this link on the map →saved by
related reading
- [1409.3215] Sequence to Sequence Learning with Neural Networksarxiv.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- [2501.00663] Titans: Learning to Memorize at Test Timearxiv.org
- Understanding Memorization via Loss Curvaturegoodfire.ai
- The World Inside Neural Networksgoodfire.ai
- NL.pdfabehrouz.github.io
- [2602.15029] Symmetry in language statistics shapes the geometry of model representationsarxiv.org
- [2501.12352] Test-time regression: a unifying framework for designing sequence models with associative memoryar5iv.labs.arxiv.org
- Transformers from Scratche2eml.school
- Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1) — AI Alignment Forumalignmentforum.org
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum · GitHubgithub.com
- arxiv.org/pdf/2505.24832arxiv.org