Frontier language models have become much smaller | Epoch AI
In this Gradient Updates weekly issue, Ege discusses how frontier language models have unexpectedly reversed course on scaling, with current models an order of magnitude smaller than GPT-4.
Frontier language models have become much smaller | Epoch AI Gradient Updates shares more opinionated or informal takes on big questions in AI progress. These posts solely represent the views of the authors, and do not necessarily reflect the views of Epoch AI as a whole. Between the release of the original Transformer in 2017 and the release of GPT-4, language models at the frontier of capabilities became much larger. Parameter counts were scaled up by 1000 times from 117 million to 175 billion between GPT-1 and GPT-3 in the span of two years and by another 10 times from 175 billion to 1.8 tr
related reading
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- Scaling: The State of Play in AIoneusefulthing.org
- AI progress is about to speed up | Epoch AIepoch.ai
- Chinchillaarxiv.org
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- My picture of the present in AI — LessWronglesswrong.com
- Pre-training isn't dead, it’s just restingtmychow.substack.com
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- Fermi estimate of future training runsdanieldewey.net
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com