1810.04805
arxiv.org · 6,675 words · saved by 2 readers
N/A
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding Jacob Devlin Ming-Wei Chang Kenton Lee Kristina Toutanova Google AI Language {jacobdevlin,mingweichang,kentonl,kristout}@google.com Abstract…
saved by
related reading
- 2005.14165arxiv.org
- Generalized Language Models | Lil'Loglilianweng.github.io
- Transfer Learninglena-voita.github.io
- The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning) – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning) – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Recent Advances in Language Model Fine-tuningruder.io
- 1706.03762arxiv.org
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- radford2018improving.pdfcs.ubc.ca
- TOD-Bertarxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- [2106.09685] LoRA: Low-Rank Adaptation of Large Language Modelsarxiv.org