flâneur

[1606.08415] Gaussian Error Linear Units (GELUs)

arxiv.org · 4,248 words · saved by 1 readers

Abstract:We propose the Gaussian Error Linear Unit (GELU), a high-performing neural network activation function. The GELU activation function is $x\Phi(x)$, where $\Phi(x)$ the standard Gaussian cumulative distribution function. The GELU nonlinearity weights inputs by their value, rather than gates inputs by their sign as in ReLUs ($x\mathbf{1}_{x>0}$). We perform an empirical evaluation of the GELU nonlinearity against the ReLU and ELU activations and find performance improvements across all considered computer vision, natural language processing, and speech tasks.

G AUSSIAN E RROR L INEAR U NITS (GELU S ) Dan Hendrycks∗ Kevin Gimpel University of California, Berkeley Toyota Technological Institute at Chicago hendrycks@berkeley.edu kgimpel@ttic.edu A BSTRACT We propose the Gaussian Error Linear Unit (GELU), a…

related reading