[2206.07682] Emergent Abilities of Large Language Models
Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models. We consider an ability to be emergent if it is not present in smaller models but is present in larger models. Thus, emergent abilities cannot be predicted simply by extrapolating the performance of smaller models. The existence of such emergence implies that additional scaling could further expand the range of capabilities of language models.
Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models. We consider an ability to be emergent if it is not present in smaller models but is present in larger models. Thus, emergent abilities cannot be predicted simply by extrapolating the performance of smaller models. The existence of such emergence implies that additional scaling could further expand the range of capabilities of language model
Explore this link on the map →related reading
- Are Emergent Abilities in Large Language Models just In-Context Learning?arxiv.org
- [2309.01809] Are Emergent Abilities in Large Language Models just In-Context Learning?arxiv.org
- On neural scaling and the quanta hypothesisericjmichaud.com
- The Scaling Hypothesis · Gwern.netgwern.net
- [2405.10938] Observational Scaling Laws and the Predictability of Language Model Performancearxiv.org
- larger language models may disappoint you [or, an eternally unfinished draft] — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- Large Language Model: world models or surface statistics?thegradient.pub
- irmckenzie.co.uk/round2irmckenzie.co.uk
- Thoughts on the Alignment Implications of Scaling Language Models | Leo Gaobmk.sh
- GitHub - inverse-scaling/prize: A prize for finding tasks that cause large language models to show inverse scaling · GitHubgithub.com