✳flâneur — a map of the web's best reading
Introducing BART | TensorGoose
sshleifer.github.io · 2,479 words · saved by 1 readers
Episode 1 – a mysterious new Seq2Seq model with state of the art summarization performance visits a popular open source library
Overview Background: Seq2Seq Pretraining Bert vs. GPT2 Encoder-Decoder Pretraining: Fill In the Span Summarization Demo: BartForConditionalGeneration Conclusion Overview For the past few weeks, I worked on integrating BART into transformers . This post covers the high-level differences between BART and its predecessors and how to use the new BartForConditionalGeneration to summarize documents. Leave a comment below if you have any questions! Background: Seq2Seq Pretraining In October 2019, teams from Google and Facebook published new transformer papers: T5 and BART . Both papers achieved bette
Explore this link on the map →related reading
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Transformers from Scratche2eml.school
- [1706.03762] Attention Is All You Needarxiv.org
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- Gwern visits BAIR – Yuxi on the Wiredyuxi.ml
- Seq2seq and Attentionlena-voita.github.io
- The Annotated Transformernlp.seas.harvard.edu
- Generalized Language Models | Lil'Loglilianweng.github.io
- Thread by @karpathy on Thread Reader App – Thread Reader Appthreadreaderapp.com
- Transformer Architectures · Hugging Facehuggingface.co