From Deep to Long Learning? · Hazy Research
For the last two years, a line of work in our lab has been to increase sequence length. We thought longer sequences would enable a new era of machine learning foundation models: they could learn from longer contexts, multiple media sources, complex demonstrations, and more. All data ready and waiting to be learned from in the world! It’s been amazing to see the progress there. As an aside, we’re happy to play a role with the introduction of FlashAttention (code, blog, paper) by Tri Dao and Dan Fu from our lab, who showed that sequence lengths of 32k are possible–and now widely available in this era of foundation models (and we’ve heard OpenAI, Microsoft, NVIDIA, and others use it for their models too–awesome!). The context lengths of foundation models have been growing recently! What's next? As the GPT4 press release noted, this has allowed almost 50 pages of text as context–and tokenization/patching ideas like those in Deepmind’s Gato are able to use images as context. So many amazing
From Deep to Long Learning? · Hazy Research Mar 28, 2023 · 9 min read From Deep to Long Learning? Dan Fu , Michael Poli , Chris Ré . For the last two years , a line of work in our lab has been to increase sequence length. We thought longer sequences would enable a new era of machine learning foundation models: they could learn from longer contexts, multiple media sources, complex demonstrations, and more. All data ready and waiting to be learned from in the world! It’s been amazing to see the progress there. As an aside, we’re happy to play a role with the introduction of FlashAttention ( code
Explore this link on the map →related reading
- Hyena Hierarchy: Towards Larger Convolutional Language Models · Hazy Researchhazyresearch.stanford.edu
- Can Longer Sequences Help Take the Next Leap in AI? | SAIL Blogai.stanford.edu
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- H3: Language Modeling with State Space Models and (Almost) No Attention · Hazy Researchhazyresearch.stanford.edu
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- [2111.00396] Efficiently Modeling Long Sequences with Structured State Spacesarxiv.org
- Mamba: The Easy Wayjackcook.com
- [2507.04239] Scaling Context Requires Rethinking Attentionarxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- [2212.14052] Hungry Hungry Hippos: Towards Language Modeling with State Space Modelsarxiv.org
- 1706.03762arxiv.org