flâneur

Learning to Remember — Lawrence Feng

ar-forum.github.io · 1,542 words · saved by 1 readers

Two models can look identical at the end of post-training and diverge once you fine-tune them. That is, how a capability was acquired affects how much of it survives. In our work, we leverage this perspective to identify training choices that push the retention–adaptation tradeoff outward.

COLM 2026 Early data exposure improves robustness to subsequent fine-tuning. Two models can look identical at the end of post-training and diverge once you fine-tune them. That is, how a capability was acquired affects how much of it survives. In our work, we leverage this perspective to identify training choices that push the retention–adaptation tradeoff outward. Part I · The setup Downstream forgetting is an upstream problem. When a post-trained model is released for downstream fine-tuning, its carefully acquired capabilities are at risk. Fine-tuning on a new objective routinely…

saved by

related reading