Architecture Research as Addressing Constraints to Scaling
Author’s note: Short note. Really a quick extension to scaling is subtler than it seems. This is probably obvious to many but still worth writing, I think/hope. Sometimes when I talk to people about architecture research, they have a somewhat naive impression that it is mostly about empirically searching to...
Author’s note: Short note. Really a quick extension to scaling is subtler than it seems. This is probably obvious to many but still worth writing, I think/hope. Sometimes when I talk to people about architecture research, they have a somewhat naive impression that it is mostly about empirically searching to find tweaks that lower loss/perplexity a small amount. E.g. we try adding a layernorm here or there or changing this activation function or adding a convolution layer or whatever, and just trying to improve loss. This is often coupled with the idea that LLM architecture research is…
saved by
related reading
- The Scaling Hypothesis · Gwern.netgwern.net
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com
- How To Scale Your Modeljax-ml.github.io
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- On neural scaling and the quanta hypothesisericjmichaud.com
- Scaling is subtler than it seemsberen.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Will scaling work?dwarkeshpatel.com
- I am worried about near-term non-LLM AI developments — LessWronglesswrong.com
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- A short note on some aspects of long context attention | nor's blognor-blog.pages.dev