The Bitter Lesson - by Finbarr Timbers
The Bitter Lesson is an excellent essay which is overwhelmingly misunderstood. The point of the bitter lesson is that, over time, methods which scale with compute will outperform methods that do not. It is not: The idea that we should never incorporate human knowledge The idea that deep learning and scale are all we need (Rich is actually relatively skeptical of deep learning) The entire point of the essay is that, in the last 5 decades, we have seen massive increases in the amount of compute available to us as an industry and we expect to continue to see massive increases in the amount of compute available to AI research. Methods which take advantage of compute will benefit, and those that do not will suffer. The reason the lesson is bitter is that it is often much easier and quicker to get results by incorporating human knowledge. If you’re training an autocomplete system in 1995, you’re probably not going to get very far with next token prediction, and instead, handcoded, or statist
The Bitter Lesson is an excellent essay which is overwhelmingly misunderstood. The point of the bitter lesson is that, over time, methods which scale with compute will outperform methods that do not. It is not: The idea that we should never incorporate human knowledge The idea that deep learning and scale are all we need (Rich is actually relatively skeptical of deep learning) The entire point of the essay is that, in the last 5 decades, we have seen massive increases in the amount of compute available to us as an industry and we expect to continue to see massive increases in the amount…
related reading
- Does the Bitter Lesson Have Limits?dbreunig.com
- The Bitter Lesson – Yuxi on the Wiredyuxi-liu-wired.github.io
- The Bitter Lessonincompleteideas.net
- The Bitter Lessonincompleteideas.net
- The Bitter Lesson is Misunderstoodobviouslywrong.substack.com
- The Scaling Hypothesis · Gwern.netgwern.net
- As Rocks May Think | Eric Jangevjang.com
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- Dario Amodei — "We are near the end of the exponential"dwarkesh.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com