[2304.13343] Unleashing Infinite-Length Input Capacity for Large-scale Language Models with Self-Controlled Memory System
Large-scale Language Models (LLMs) are constrained by their inability to process lengthy inputs. To address this limitation, we propose the Self-Controlled Memory (SCM) system to unleash infinite-length input capacity for large-scale language models. Our SCM system is composed of three key modules: the language model agent, the memory stream, and the memory controller. The language model agent iteratively processes ultra-long inputs and stores all historical information in the memory stream. The memory controller provides the agent with both long-term memory (archived memory) and short-term memory (flash memory) to generate precise and coherent responses. The controller determines which memories from archived memory should be activated and how to incorporate them into the model input. Our SCM system can be integrated with any LLMs to enable them to process ultra-long texts without any modification or fine-tuning. Experimental results show that our SCM system enables LLMs, which are not optimized for multi-turn dialogue, to achieve multi-turn dialogue capabilities that are comparable to ChatGPT, and to outperform ChatGPT in scenarios involving ultra-long document summarization or long-term conversations. Additionally, we will supply a test set, which covers common long-text input scenarios, for evaluating the abilities of LLMs in processing long documents.~\footnote{Working in progress.}\footnote{\url{this https URL}}
Enhancing Large Language Model with Self-Controlled Memory Framework Bing Wang1 , Xinnian Liang1 , Jian Yang1 , Hui Huang2 Shuangzhi Wu3 , Peihao Wu3 , Lu Lu3 , Zejun Ma3 Zhoujun Li1 1 State Key Lab of Software Development Environment, Beihang University, Beijing, China…
saved by
related reading
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- [2603.23516] MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokensarxiv.org
- [2605.15156] MeMo: Memory as a Modelarxiv.org
- Productizing Large Language Modelsblog.replit.com
- Reimagining LLM Memory: Using Context as Training Data Unlocks Models That Learn at Test-Time | NVIDIA Technical Blogdeveloper.nvidia.com
- [2506.06266] Cartridges: Lightweight and general-purpose long context representations via self-studyarxiv.org
- Large Language Diffusion Modelsarxiv.org
- [2501.00663] Titans: Learning to Memorize at Test Timearxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Mass-Editing Memory in a Transformerarxiv.org
- Alex L. Zhangalexzhang13.github.io
- A History of Large Language Modelsgregorygundersen.com