H-Nets - the Past | Goomba Lab
goombalab.github.io · 5,852 words · saved by 6 readers
Homepage of the Goomba AI Lab @ CMU MLD.
H-Nets - the Past | Goomba Lab H-Nets - the Past This post is part of a two-part series. H-Nets: the Past H-Nets: the Future [ Paper ] [ Code ] The H-Net model has been a dream of mine for years. I’ve been fortunate enough to be able to work with my student Sukjun Hwang who made the technical breakthroughs behind this model, solving (or at least making serious progress towards) what I consider to be a very difficult but foundational problem for deep learning. In this post, I provide an informal personal recounting of the motivation and history of this project, adding various context and discus
saved by
related reading
- Dynamic Chunking for End-to-End Hierarchical Sequence Modelingarxiv.org
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- H-Nets - the Future | Goomba Labgoombalab.github.io
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- H3: Language Modeling with State Space Models and (Almost) No Attention · Hazy Researchhazyresearch.stanford.edu
- dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learningarxiv.org
- MambaByte: Token-free Selective State Space Modelarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- 2312.00752arxiv.org