flâneur — a map of the web's best reading

Review for NeurIPS paper: Language Models are Few-Shot Learners

proceedings.neurips.cc · 2,587 words · saved by 1 readers

Summary and Contributions: The paper introduces GPT-3, a very large-scale Transformer language model of 175B parameters trained on 400B tokens from CommonCrawl data. The model obtains surprisingly effective results on zero-shot and few-shot scenario, without any finetuning. With only a prompt, or conditioning on a few examples, GPT-3 obtains strong performance on a wide variety of tasks, showing that large-scale language models, while only accessing isolated text data without any other modality and being trained on a very simple task, builds an impressive understanding of human natural language. I believe this work constitutes a seminal paper that significantly advances our understanding of what's possible in natural language understanding. It is one of the most interesting paper I have read in the deep learning era. I would fight for this paper to be accepted as an oral presentation at NeurIPS 2020 and I further recommend it for the best paper award. Strengths: The paper in one of the

Review for NeurIPS paper: Language Models are Few-Shot Learners NeurIPS 2020 Language Models are Few-Shot Learners Review 1 Summary and Contributions : The paper introduces GPT-3, a very large-scale Transformer language model of 175B parameters trained on 400B tokens from CommonCrawl data. The model obtains surprisingly effective results on zero-shot and few-shot scenario, without any finetuning. With only a prompt, or conditioning on a few examples, GPT-3 obtains strong performance on a wide variety of tasks, showing that large-scale language models, while only accessing isolated text data wi

Explore this link on the map →

related reading