Extending AFM-4.5B to 64k Context Length
From 4k to 64k context through aggressive experimentation, model merging, distillation, and a concerning amount of soup. The other day Arcee finally announced the first of our from-scratch foundation models, AFM-4.5B. Learning to train a foundation model is a long and arduous journey, and there are many lessons and learnings that we will be sharing in the full tech report in the coming weeks. In the meantime, I wanted to pull back the curtain on one particular part of the training process: extending the context length. We extended AFM-4.5B from 4k to 64k context through aggressive experimentation, model merging, distillation, and a concerning amount of soup. This post will be a pretty unflattering look at the raw meat of the experimental process and the various approaches we tried, eventually arriving at a final model that performs well on both short and long context tasks. Bon appétit. Disclaimer: AFM-4.5B was recently introduced as a preview, with a full open-weight release (under a
Arcee AI | Extending AFM-4.5B to 64k Context Length Trinity Large Thinking: Available on OpenRouter. Try now ↗ ENTERPRISE Research COMPANY Get API Blog / Extending AFM-4.5B to 64k Context Length Extending AFM-4.5B to 64k Context Length Charles Goddard , • June 23, 2025 From 4k to 64k context through aggressive experimentation, model merging, distillation, and a concerning amount of soup. The other day Arcee finally announced the first of our from-scratch foundation models, AFM-4.5B . Learning to train a foundation model is a long and arduous journey, and there are many lessons and learnings th
Explore this link on the map →saved by
related reading
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- Extending the Context of Pretrained LLMs by Dropping their Positional Embeddingspub.sakana.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Reimagining LLM Memory: Using Context as Training Data Unlocks Models That Learn at Test-Time | NVIDIA Technical Blogdeveloper.nvidia.com
- Composer2.pdfcursor.com
- From Deep to Long Learning? · Hazy Researchhazyresearch.stanford.edu
- Generalizing an LLM from 8k to 1M Context using Qwen-Agent | Qwenqwenlm.github.io
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev
- Subquadratic — How SSA Makes Long Context Practicalsubq.ai
- gemini_v1_5_report.pdfstorage.googleapis.com
- Why We Need Continual Learning | Andreessen Horowitza16z.com