flâneur

🧠 MLPs are Hebbian Memories: A Simple Recipe for Fact-Storing Transformers · Hazy Research

hazyresearch.stanford.edu · 1,154 words · saved by 1 readers

A kernel-memory view of fact-storing MLPs gives a closed-form construction, optimal capacity guarantees, and the first theoretical account of fact storage inside Transformer blocks.

⚡ TL;DR It sounds like sci-fi, but it's real: we can instantly build knowledge into a Transformer block, no training required! Previously, we showed that MLPs can be constructed to store facts. Since then, we found a much simpler and more powerful explanation: a Transformer MLP is naturally a Hebbian memory. This view lets us develop the first closed-form fact-storing MLP construction, requiring no gradient descent 🤯, that packs facts at the information-theoretically optimal rate and is usable within Transformer blocks for factual recall! Full paper team: Roberto Garcia*, Jerry Liu*,…

saved by

related reading