flâneur — a map of the web's best reading

[1606.08415] Gaussian Error Linear Units (GELUs)

arxiv.org · 612 words · saved by 1 readers

Abstract:We propose the Gaussian Error Linear Unit (GELU), a high-performing neural network activation function. The GELU activation function is $x\Phi(x)$, where $\Phi(x)$ the standard Gaussian cumulative distribution function. The GELU nonlinearity weights inputs by their value, rather than gates inputs by their sign as in ReLUs ($x\mathbf{1}_{x>0}$). We perform an empirical evaluation of the GELU nonlinearity against the ReLU and ELU activations and find performance improvements across all considered computer vision, natural language processing, and speech tasks.

[1606.08415] Gaussian Error Linear Units (GELUs) Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:1606.08415 (cs) [Submitted on 27 Jun 2016 ( v1 ), last revised 6 Jun 2023 (this version, v5)] Title: Gaussian Error Linear Units (GELUs) Authors: Dan Hendrycks , Kevin Gimpel View a PDF of the paper titled Gaussian Error Linear Units (GELUs), by Dan Hendrycks and Kevin Gimpel View PDF Abstract: We propose the Gaussian Error Linear Unit (GELU), a high-performing ne

Explore this link on the map →

related reading