Norman Mu | The Myth of Data Inefficiency in Large Language Models
A popular assertion among cognitive psychologists is that large language models (LLMs) rely on a “developmentally implausible” form of language learning. The...
A popular assertion among cognitive psychologists is that large language models (LLMs) rely on a “developmentally implausible” form of language learning. There are many points of discordance between AI and human development, but data inefficiency is one of the most commonly discussed. Specifically, it has been claimed that LLMs “receive around four or five orders of magnitude more language data than human children” (Frank, 2023). In concrete terms: the 7 billion parameter Llama-2 model was pre-trained on 2 trillion tokens, and the 100 million token BabyLM dataset proposed by Warstadt et al. (2
related reading
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Will scaling work?dwarkeshpatel.com
- How can LLM RL Work Despite Information-Theoretic Inefficiencyberen.io
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- The real data wall is billions of years of evolutiondynomight.net
- chinchilla's wild implications — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- Headroom for AI development – Machine Learning (Theory)hunch.net
- Scaling: The State of Play in AIoneusefulthing.org
- Do large language models understand us? | by Blaise Aguera y Arcas | Mediummedium.com
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org