Five years of GPT progress
In this article, I discuss the generative pre-trained transformer (GPT) line of work, and how it has evolved over time. I focus on the SOTA models, and the differences between them. There are a bunch of different articles summarizing these papers, but nothing that I’m aware of that explicitly focuses on the differences between them. I focus on the GPT line of research as that’s what’s driving the current fever pitch of development. There’s a ton of prior work before large GPTs (eg the n-gram models from the 2000s, BERT, etc) but this post is super long, so I’m gonna save those for future articles. Abstract The first GPT paper is interesting to read in hindsight. It doesn’t appear like anything special and doesn’t follow any of the conventions that have developed. The dataset is described in terms of GB rather than tokens, and the number of parameters in the model isn’t explicitly stated. To a certain extent, I suspect that the paper was a side project at OpenAI and wasn’t viewed as par
Five years of GPT progress Finbarr Timbers Blog Books Advice LLM Performance Tools Five years of GPT progress If you want to read more of my writing, I have a Substack . In this article, I discuss the generative pre-trained transformer (GPT) line of work, and how it has evolved over time. I focus on the SOTA models, and the differences between them. There are a bunch of different articles summarizing these papers, but nothing that I’m aware of that explicitly focuses on the differences between them. I focus on the GPT line of research as that’s what’s driving the current fever pitch of develop
Explore this link on the map →related reading
- How To Scale Your Modeljax-ml.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- Google "We Have No Moat, And Neither Does OpenAI"semianalysis.com
- GPT in 60 Lines of NumPy | Jay Modyjaykmody.com
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- microgptkarpathy.github.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Understanding Large Language Modelsmagazine.sebastianraschka.com
- GPT-3 - Wikipediaen.wikipedia.org
- How does GPT-3 spend its 175B parameters? — LessWronglesswrong.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com