✳flâneur — a map of the web's best reading
What happened to BERT & T5? On Transformer Encoders, PrefixLM and Denoising Objectives — Yi Tay
yitay.net · 2,414 words · saved by 1 readers
A Blogpost series about Model Architectures Part 1: What happened to BERT and T5? Thoughts on Transformer Encoders, PrefixLM and Denoising objectives
What happened to BERT & T5? On Transformer Encoders, PrefixLM and Denoising Objectives Jul 16 Written By Yi Tay The people who worked on language and NLP about 5+ years ago are left scratching their heads about where all the encoder models went. If BERT worked so well, why not scale it? What happened to encoder-decoders or encoder-only models? Today I try to unpack all that is going on, in this new era of LLMs. I Hope this post will be helpful. Few months ago I was also writing a long tweet-reply to this tweet by @srush_nlp at some point. Then the tweet got deleted because I closed the tab by
Explore this link on the map →saved by
related reading
- Transformer Architectures · Hugging Facehuggingface.co
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- transformer_attention.pdfarxiv.org
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Generalized Language Models | Lil'Loglilianweng.github.io
- The Annotated Transformernlp.seas.harvard.edu
- How do Transformers work? · Hugging Facehuggingface.co
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- The Annotated Transformernlp.seas.harvard.edu
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- Seq2seq and Attentionlena-voita.github.io