✳flâneur — a map of the web's best reading
Seq2seq and Attention
lena-voita.github.io · 6,755 words · saved by 1 readers
Sequence to sequence models (training and inference), the concept of attention and the Transformer model.
Seq2seq and Attention p { text-align: justify; } ⇤ NLP Course | For You Seq2seq and Attention Seq2seq Basics • Intro • Encoder-Decoder Framework • Conditional LMs • The Simplest Model: RNNs • Training • Inference: Beam Search Attention • Why do we need it? • Attention: High-Level • Attention Score Functions • Models: Bahdanau vs Luong • Attention and Alignment Transformer • Intro • Self-Attention • Masked Self-Attention • Multi-Head Attention • Model Architecture Subword Segmentation: BPE Analysis a
Explore this link on the map →saved by
related reading
- Transformers Explained Visually (Part 1): Overview of Functionality | Towards Data Sciencetowardsdatascience.com
- transformer_attention.pdfarxiv.org
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- 1706.03762arxiv.org
- The Illustrated Transformer – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Visualizing A Neural Machine Translation Model (Mechanics of Seq2seq Models With Attention) – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- Transformers from Scratche2eml.school
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Everything About Transformerskrupadave.com
- [1706.03762] Attention Is All You Needarxiv.org
- Attention is all you need: Discovering the Transformer paper | Towards Data Sciencetowardsdatascience.com
- Transformer (deep learning) - Wikipediaen.wikipedia.org