Norman Mu | The Myth of Data Inefficiency in Large Language Models
A popular assertion among cognitive psychologists is that large language models (LLMs) rely on a “developmentally implausible” form of language learning. The...
A popular assertion among cognitive psychologists is that large language models (LLMs) rely on a “developmentally implausible” form of language learning. There are many points of discordance between AI and human development, but data inefficiency is one of the most commonly discussed. Specifically, it has been claimed that LLMs “receive around four or five orders of magnitude more language data than human children” (Frank, 2023). In concrete terms: the 7 billion parameter Llama-2 model was pre-trained on 2 trillion tokens, and the 100 million token BabyLM dataset proposed by Warstadt et al. (2
Explore this link on the map →related reading
- There Are No New Ideas in AI… Only New Datasetsblog.jxmo.io
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- chinchilla's wild implications — LessWronglesswrong.com
- Headroom for AI development – Machine Learning (Theory)hunch.net
- GenAI Handbookgenai-handbook.github.io
- Will scaling work? - by Dwarkesh Patel - Dwarkesh Podcastdwarkesh.com
- chinchilla's wild implications — AI Alignment Forumalignmentforum.org
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- Do large language models understand us? | by Blaise Aguera y Arcas | Mediummedium.com
- Language Models in Plato's Cave - by Sergey Levinesergeylevine.substack.com
- Large Language Model: world models or surface statistics?thegradient.pub
- Fermi estimate of future training runsdanieldewey.net