flâneur — a map of the web's best reading

Encoder-Only vs Decoder-Only vs Encoder-Decoder Transformer

vaclavkosar.com · 929 words · saved by 1 readers

People keep asking me about, what is the difference between encoder, decoder, and normal transformer (with self-attention). It is a simple thing, you can master quickly. BERT has Encoder-only architecture. Input is text and output is sequence of embeddings. Use cases are sequence classification (class token), token classification. It uses bidirectional attention, so the model can see forwards and backwards. Another encoder-only model example is ViT (Vision Transformer) for image classification. GPT-2 has Decoder-only architecture. Input is text and output is the next word (token), which is then appended to the input. Use cases are mostly text generation (autoregressive), but with prompting we can do many things including sequence classification. The attention is almost always causal (unidirectional), so the model can see only previous tokens (prefix). T5 has Encoder-Decoder or Full-Transformer. Input is text and output is the next word (token), which is then appended to the decoder-inp

PG138 Login dan Daftar Akun Resmi Melalui Autentikasi Keamanan Siber LOGIN DAFTAR PG138 Login dan Daftar Akun Resmi Melalui Autentikasi Keamanan Siber Log In | Sign Up PG138 AGEN 338 RTP PG138 SITUS PG138 PG138 RTP LIVE LINK ALTERNATIF PG138 Search SUGGESTIONS PRODUCTS Search for RECENT SEARCH Hapus riwayat YOU MAY ALSO LIKE Mid Autumn of Love Mooncake Rp 588.000 (-40%) Rp 348.000 Mid Autumn of Joy Mooncake Rp 888.000 (-33%) Rp 588.000 Mid Autumn of Fortune Mooncake Rp 1.188.000 (-33%) Rp 788.000 Mid Autumn Festival Mooncake Rp 1.288.000 (-23%) Rp 988.000 The Jade Bouquet Rp 435.000 (-34%) Rp

Explore this link on the map →

related reading