Extending Context is Hard | kaiokendev.github.io
kaiokendev.github.io · 8,182 words · saved by 3 readers
pages
Extending Context is Hard | kaiokendev kaiokendev Extending Context is Hard…but not Impossible† On the surface, it should be an easy task. I was working on this write up while I was working through methods to finetune the pre-trained model for longer sequence length. In this case, the pre-trained model is LLaMa, with a pre-training sequence length of 2048. Naively fine-tuning the model on long sequences never seemed to work, but I felt it must be possible, so I stubbornly pushed through it. Now, there is a way to extend context with just 1 line of code, and it is getting a lot of attention. Un
saved by
related reading
- [2108.12409] Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolationarxiv.org
- Extending the Context of Pretrained LLMs by Dropping their Positional Embeddingspub.sakana.ai
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolationarxiv.org
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev
- [2512.23675] End-to-End Test-Time Training for Long Contextarxiv.org
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- [2507.04239] Scaling Context Requires Rethinking Attentionarxiv.org
- GLM-5.2: Built for Long-Horizon Tasksz.ai
- Arcee AI | Extending AFM-4.5B to 64k Context Lengtharcee.ai