Learning to Remember — Lawrence Feng
Two models can look identical at the end of post-training and diverge once you fine-tune them. That is, how a capability was acquired affects how much of it survives. In our work, we leverage this perspective to identify training choices that push the retention–adaptation tradeoff outward.
COLM 2026 Early data exposure improves robustness to subsequent fine-tuning. Two models can look identical at the end of post-training and diverge once you fine-tune them. That is, how a capability was acquired affects how much of it survives. In our work, we leverage this perspective to identify training choices that push the retention–adaptation tradeoff outward. Part I · The setup Downstream forgetting is an upstream problem. When a post-trained model is released for downstream fine-tuning, its carefully acquired capabilities are at risk. Fine-tuning on a new objective routinely…
saved by
related reading
- Early Data Exposure Improves Robustness to Subsequent Fine-Tuningarxiv.org
- Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggersarxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Continual Learning Problemjessylin.com
- Recent Advances in Language Model Fine-tuningruder.io
- Modern Pretraining Strategies: A Hands-On Guidetheneuralmaze.substack.com
- Understanding Memorization via Loss Curvaturegoodfire.ai
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- RL's Razor: Why Online Reinforcement Learning Forgets Lessarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- [1606.09282] Learning without Forgettingarxiv.org
- Shaping capabilities with token-level data filteringarxiv.org