✳flâneur — a map of the web's best reading
Extending Context is Hard | kaiokendev.github.io
kaiokendev.github.io · 8,182 words · saved by 2 readers
pages
Extending Context is Hard | kaiokendev kaiokendev Extending Context is Hard…but not Impossible† On the surface, it should be an easy task. I was working on this write up while I was working through methods to finetune the pre-trained model for longer sequence length. In this case, the pre-trained model is LLaMa, with a pre-training sequence length of 2048. Naively fine-tuning the model on long sequences never seemed to work, but I felt it must be possible, so I stubbornly pushed through it. Now, there is a way to extend context with just 1 line of code, and it is getting a lot of attention. Un
Explore this link on the map →saved by
related reading
- Extending the Context of Pretrained LLMs by Dropping their Positional Embeddingspub.sakana.ai
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Arcee AI | Extending AFM-4.5B to 64k Context Lengtharcee.ai
- Reimagining LLM Memory: Using Context as Training Data Unlocks Models That Learn at Test-Time | NVIDIA Technical Blogdeveloper.nvidia.com
- Mediumblog.gopenai.com
- [2507.04239] Scaling Context Requires Rethinking Attentionarxiv.org
- The Annotated Transformernlp.seas.harvard.edu