Frontier language models have become much smaller | Epoch AI
In this Gradient Updates weekly issue, Ege discusses how frontier language models have unexpectedly reversed course on scaling, with current models an order of magnitude smaller than GPT-4.
Frontier language models have become much smaller | Epoch AI Gradient Updates shares more opinionated or informal takes on big questions in AI progress. These posts solely represent the views of the authors, and do not necessarily reflect the views of Epoch AI as a whole. Between the release of the original Transformer in 2017 and the release of GPT-4, language models at the frontier of capabilities became much larger. Parameter counts were scaled up by 1000 times from 117 million to 175 billion between GPT-1 and GPT-3 in the span of two years and by another 10 times from 175 billion to 1.8 tr
Explore this link on the map →related reading
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- AI progress is about to speed up | Epoch AIepoch.ai
- My picture of the present in AI — LessWronglesswrong.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Demystify Transformers: A Guide to Scaling Laws | by Yu-Cheng Tsai | Sage Ai | Mediummedium.com
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- Fermi estimate of future training runsdanieldewey.net
- State of AI 2025: 100T Token LLM Usage Study | OpenRouteropenrouter.ai
- things that confuse me about the current AI market. — LessWronglesswrong.com
- GPT-4 Architecture, Infrastructure, Training Dataset, Costs, Vision, MoEsemianalysis.com