Review for NeurIPS paper: Language Models are Few-Shot Learners
Summary and Contributions: The paper introduces GPT-3, a very large-scale Transformer language model of 175B parameters trained on 400B tokens from CommonCrawl data. The model obtains surprisingly effective results on zero-shot and few-shot scenario, without any finetuning. With only a prompt, or conditioning on a few examples, GPT-3 obtains strong performance on a wide variety of tasks, showing that large-scale language models, while only accessing isolated text data without any other modality and being trained on a very simple task, builds an impressive understanding of human natural language. I believe this work constitutes a seminal paper that significantly advances our understanding of what's possible in natural language understanding. It is one of the most interesting paper I have read in the deep learning era. I would fight for this paper to be accepted as an oral presentation at NeurIPS 2020 and I further recommend it for the best paper award. Strengths: The paper in one of the
Review for NeurIPS paper: Language Models are Few-Shot Learners NeurIPS 2020 Language Models are Few-Shot Learners Review 1 Summary and Contributions : The paper introduces GPT-3, a very large-scale Transformer language model of 175B parameters trained on 400B tokens from CommonCrawl data. The model obtains surprisingly effective results on zero-shot and few-shot scenario, without any finetuning. With only a prompt, or conditioning on a few examples, GPT-3 obtains strong performance on a wide variety of tasks, showing that large-scale language models, while only accessing isolated text data wi
Explore this link on the map →related reading
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- The Scaling Hypothesis · Gwern.netgwern.net
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- gpt-4.pdfcdn.openai.com
- More Efficient In-Context Learning with GLaMblog.research.google
- larger language models may disappoint you [or, an eternally unfinished draft] — LessWronglesswrong.com
- Understanding Large Language Modelsmagazine.sebastianraschka.com
- Generalized Language Models | Lil'Loglilianweng.github.io
- Recent Advances in Language Model Fine-tuningruder.io
- GPT-3 - Wikipediaen.wikipedia.org