[1606.08415] Gaussian Error Linear Units (GELUs)
Abstract:We propose the Gaussian Error Linear Unit (GELU), a high-performing neural network activation function. The GELU activation function is $x\Phi(x)$, where $\Phi(x)$ the standard Gaussian cumulative distribution function. The GELU nonlinearity weights inputs by their value, rather than gates inputs by their sign as in ReLUs ($x\mathbf{1}_{x>0}$). We perform an empirical evaluation of the GELU nonlinearity against the ReLU and ELU activations and find performance improvements across all considered computer vision, natural language processing, and speech tasks.
[1606.08415] Gaussian Error Linear Units (GELUs) Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:1606.08415 (cs) [Submitted on 27 Jun 2016 ( v1 ), last revised 6 Jun 2023 (this version, v5)] Title: Gaussian Error Linear Units (GELUs) Authors: Dan Hendrycks , Kevin Gimpel View a PDF of the paper titled Gaussian Error Linear Units (GELUs), by Dan Hendrycks and Kevin Gimpel View PDF Abstract: We propose the Gaussian Error Linear Unit (GELU), a high-performing ne
Explore this link on the map →related reading
- [1606.08415] Gaussian Error Linear Units (GELUs)arxiv.org
- Activation function and GLU variants for Transformer models | by Tarique Anwar | Mediummedium.com
- ReLU Activation Function in Deep Learning - GeeksforGeeksgeeksforgeeks.org
- 2002.05202arxiv.org
- [2002.05202] GLU Variants Improve Transformerar5iv.labs.arxiv.org
- [2002.05202] GLU Variants Improve Transformerarxiv.org
- The Decade of Deep Learning | Leo Gaobmk.sh
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Tiny TPUtinytpu.com
- gpt-4.pdfcdn.openai.com
- Capítulo 8 - Função de Ativação - Deep Learning Bookdeeplearningbook.com.br