[2210.07229] Mass-Editing Memory in a Transformer
Abstract:Recent work has shown exciting promise in updating large language models with new memories, so as to replace obsolete information or add specialized knowledge. However, this line of work is predominantly limited to updating single associations. We develop MEMIT, a method for directly updating a language model with many memories, demonstrating experimentally that it can scale up to thousands of associations for GPT-J (6B) and GPT-NeoX (20B), exceeding prior work by orders of magnitude. Our code and data are at this https URL.
View PDF HTML (experimental) Abstract:Recent work has shown exciting promise in updating large language models with new memories, so as to replace obsolete information or add specialized knowledge. However, this line of work is predominantly limited to updating single associations. We develop MEMIT, a method for directly updating a language model with many memories, demonstrating experimentally that it can scale up to thousands of associations for GPT-J (6B) and GPT-NeoX (20B), exceeding prior work by orders of magnitude. Our code and data are at this https URL. Comments: 18 pages, 11…
saved by
related reading
- [2605.15156] MeMo: Memory as a Modelarxiv.org
- The Continual Learning Problemjessylin.com
- Understanding Memorization via Loss Curvaturegoodfire.ai
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- NL.pdfabehrouz.github.io
- [2501.00663] Titans: Learning to Memorize at Test Timearxiv.org
- Self-Adapting Language Modelsarxiv.org
- Do Language Models Need Sleep? Offline Recurrence for Improved Online Inferencearxiv.org
- arxiv.org/pdf/2505.24832arxiv.org
- MemPrompt: Memory-assisted Prompt Editing with User Feedbackmemprompt.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com