flâneur — a map of the web's best reading

Adam Optimizer Causes Privileged Basis in Transformer Language Models — LessWrong

lesswrong.com · 1,535 words · saved by 1 readers

Diego Caples (diego@activated-ai.com) • Rob Neuhaus (rob@activated-ai.com) …

x Adam Optimizer Causes Privileged Basis in Transformer LM Residual Stream — LessWrong Interpretability (ML & AI) Machine Learning (ML) Optimization Transformers AI Frontpage 74 Adam Optimizer Causes Privileged Basis in Transformer LM Residual Stream by Diego Caples , rrenaud 6th Sep 2024 4 min read 8 74 Diego Caples ( diego@activated-ai.com ) Rob Neuhaus ( rob@activated-ai.com ) Introduction In principle, neuron activations in a transformer-based language model residual stream should be about the same scale. In practice, the dimensions unexpectedly widely vary in scale. Mathematical theories

Explore this link on the map →

related reading