[2604.11811] M$^\star$: Every Task Deserves Its Own Memory Harness
Abstract:Large language model agents rely on specialized memory systems to accumulate and reuse knowledge during extended interactions. Recent architectures typically adopt a fixed memory design tailored to specific domains, such as semantic retrieval for conversations or skills reused for coding. However, a memory system optimized for one purpose frequently fails to transfer to others. To address this limitation, we introduce M$^\star$, a method that automatically discovers task-optimized memory harnesses through executable program evolution. Specifically, M$^\star$ models an agent memory system as a memory program written in Python. This program encapsulates the data Schema, the storage Logic, and the agent workflow Instructions. We optimize these components jointly using a reflective code evolution method; this approach employs a population-based search strategy and analyzes evaluation failures to iteratively refine the candidate programs. We evaluate M$^\star$ on four distinct benchmarks spanning conversation, embodied planning, and expert reasoning. Our results demonstrate that M$^\star$ improves performance over existing fixed-memory baselines robustly across all evaluated tasks. Furthermore, the evolved memory programs exhibit structurally distinct processing mechanisms for each domain. This finding indicates that specializing the memory mechanism for a given task explores a broad design space and provides a superior solution compared to general-purpose memory paradigms.
[2604.11811] M$^\star$: Every Task Deserves Its Own Memory Harness --> Computer Science > Programming Languages arXiv:2604.11811 (cs) [Submitted on 10 Apr 2026 ( v1 ), last revised 23 May 2026 (this version, v2)] Title: M$^\star$: Every Task Deserves Its Own Memory Harness Authors: Wenbo Pan , Shujie Liu , Xiangyang Zhou , Shiwei Zhang , Wanlu Shi , Mirror Xu , Xiaohua Jia View a PDF of the paper titled M$^\star$: Every Task Deserves Its Own Memory Harness, by Wenbo Pan and 6 other authors View PDF HTML (experimental) Abstract: Large language model agents rely on specialized memory systems to
Explore this link on the map →saved by
related reading
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- [2605.15156] MeMo: Memory as a Modelarxiv.org
- Composer2.pdfcursor.com
- Making Sense of Memory in AI Agents – Leonie Monigattileoniemonigatti.com
- How AI Agents Remember Thingsdamiangalarza.com
- The Continual Learning Problemjessylin.com
- Explore | alphaXivalphaxiv.org
- Language Models can Solve Computer Tasksarxiv.org
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- [2507.19457] GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learningarxiv.org
- Introducing Context Repositories: Git-based Memory for Coding Agents | Lettaletta.com
- NL.pdfabehrouz.github.io