flâneur — a map of the web's best reading

Extending AFM-4.5B to 64k Context Length

arcee.ai · 2,929 words · saved by 1 readers

From 4k to 64k context through aggressive experimentation, model merging, distillation, and a concerning amount of soup. The other day Arcee finally announced the first of our from-scratch foundation models, AFM-4.5B. Learning to train a foundation model is a long and arduous journey, and there are many lessons and learnings that we will be sharing in the full tech report in the coming weeks. In the meantime, I wanted to pull back the curtain on one particular part of the training process: extending the context length. We extended AFM-4.5B from 4k to 64k context through aggressive experimentation, model merging, distillation, and a concerning amount of soup. This post will be a pretty unflattering look at the raw meat of the experimental process and the various approaches we tried, eventually arriving at a final model that performs well on both short and long context tasks. Bon appétit. Disclaimer: AFM-4.5B was recently introduced as a preview, with a full open-weight release (under a

Arcee AI | Extending AFM-4.5B to 64k Context Length Trinity Large Thinking: Available on OpenRouter. Try now ↗ ENTERPRISE Research COMPANY Get API Blog / Extending AFM-4.5B to 64k Context Length Extending AFM-4.5B to 64k Context Length Charles Goddard , • June 23, 2025 From 4k to 64k context through aggressive experimentation, model merging, distillation, and a concerning amount of soup. The other day Arcee finally announced the first of our from-scratch foundation models, AFM-4.5B . Learning to train a foundation model is a long and arduous journey, and there are many lessons and learnings th

Explore this link on the map →

saved by

related reading