🧠 MLPs are Hebbian Memories: A Simple Recipe for Fact-Storing Transformers · Hazy Research
A kernel-memory view of fact-storing MLPs gives a closed-form construction, optimal capacity guarantees, and the first theoretical account of fact storage inside Transformer blocks.
⚡ TL;DR It sounds like sci-fi, but it's real: we can instantly build knowledge into a Transformer block, no training required! Previously, we showed that MLPs can be constructed to store facts. Since then, we found a much simpler and more powerful explanation: a Transformer MLP is naturally a Hebbian memory. This view lets us develop the first closed-form fact-storing MLP construction, requiring no gradient descent 🤯, that packs facts at the information-theoretically optimal rate and is usable within Transformer blocks for factual recall! Full paper team: Roberto Garcia*, Jerry Liu*,…
saved by
related reading
- Transformers Provably Learn to Internalize Chain-of-Thoughtarxiv.org
- How LLMs Actually Work | 0xkato0xkato.xyz
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- ali (@waterloo_intern) on Xx.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1) — AI Alignment Forumalignmentforum.org
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- NL.pdfabehrouz.github.io
- [2501.00663] Titans: Learning to Memorize at Test Timearxiv.org
- Can a Language Model Learn Facts Continually in Its Weights?labs.baseten.co
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- [2510.26745] Deep sequence models tend to memorize geometrically; it is unclear whyarxiv.org