Activation function and GLU variants for Transformer models | by Tarique Anwar | Medium
medium.com · 1,681 words · saved by 1 readers
Characterizing the first week of April 2022 as happening in the field of AI and Deep Learning would be an understatement. Within the same…
Activation function and GLU variants for Transformer models Tarique Anwar 8 min read · Apr 18, 2022 -- 1 Listen Share Characterizing the first week of April 2022 as happening in the field of AI and Deep Learning would be an understatement. Within the same week, Google and OpenAI showcased their models PaLM⁶ and DALLE 2. PaLM is a 540 billion parameter transformer-based language model seemingly capable of a state-of-the-art performance on a multitude of tasks in natural language. In addition to those, breakthrough capabilities are also demonstrated in reasoning tasks. DALLE 2 is an AI model whi
saved by
related reading
- [2002.05202] GLU Variants Improve Transformerar5iv.labs.arxiv.org
- [2002.05202] GLU Variants Improve Transformerarxiv.org
- 2002.05202arxiv.org
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- [1606.08415] Gaussian Error Linear Units (GELUs)arxiv.org
- ReLU Activation Function in Deep Learning - GeeksforGeeksgeeksforgeeks.org
- Transformers from Scratche2eml.school
- transformer_attention.pdfarxiv.org
- Transformer (deep learning) - Wikipediaen.wikipedia.org
- The Annotated Transformernlp.seas.harvard.edu
- The Annotated Transformernlp.seas.harvard.edu
- [1606.08415] Gaussian Error Linear Units (GELUs)arxiv.org