The Bitter Lesson - by Finbarr Timbers
The Bitter Lesson is an excellent essay which is overwhelmingly misunderstood. The point of the bitter lesson is that, over time, methods which scale with compute will outperform methods that do not. It is not: The idea that we should never incorporate human knowledge The idea that deep learning and scale are all we need (Rich is actually relatively skeptical of deep learning) The entire point of the essay is that, in the last 5 decades, we have seen massive increases in the amount of compute available to us as an industry and we expect to continue to see massive increases in the amount of compute available to AI research. Methods which take advantage of compute will benefit, and those that do not will suffer. The reason the lesson is bitter is that it is often much easier and quicker to get results by incorporating human knowledge. If you’re training an autocomplete system in 1995, you’re probably not going to get very far with next token prediction, and instead, handcoded, or statist
The Bitter Lesson Far too many people misunderstand the bitter lesson Finbarr Timbers Jun 26, 2025 44 7 5 Share The Bitter Lesson is an excellent essay which is overwhelmingly misunderstood. The point of the bitter lesson is that, over time, methods which scale with compute will outperform methods that do not. It is not: The idea that we should never incorporate human knowledge The idea that deep learning and scale are all we need (Rich is actually relatively skeptical of deep learning) The entire point of the essay is that, in the last 5 decades, we have seen massive increases in the amount o
Explore this link on the map →related reading
- Does the Bitter Lesson Have Limits?dbreunig.com
- The Bitter Lesson – Yuxi on the Wiredyuxi-liu-wired.github.io
- The Bitter Lessonincompleteideas.net
- The Bitter Lessonincompleteideas.net
- The Scaling Hypothesis · Gwern.netgwern.net
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- Why AGI Will Not Happen - Tim Dettmerstimdettmers.com
- Algorithmic Improvement Is Probably Faster Than Scaling Now — LessWronglesswrong.com