Adam Optimizer Causes Privileged Basis in Transformer Language Models — LessWrong
lesswrong.com · 1,535 words · saved by 1 readers
Diego Caples (diego@activated-ai.com) • Rob Neuhaus (rob@activated-ai.com) …
x Adam Optimizer Causes Privileged Basis in Transformer LM Residual Stream — LessWrong Interpretability (ML & AI) Machine Learning (ML) Optimization Transformers AI Frontpage 74 Adam Optimizer Causes Privileged Basis in Transformer LM Residual Stream by Diego Caples , rrenaud 6th Sep 2024 4 min read 8 74 Diego Caples ( diego@activated-ai.com ) Rob Neuhaus ( rob@activated-ai.com ) Introduction In principle, neuron activations in a transformer-based language model residual stream should be about the same scale. In practice, the dimensions unexpectedly widely vary in scale. Mathematical theories
related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Muon Outperforms Adam in Tail-End Associative Memory Learningarxiv.org
- The Annotated Transformernlp.seas.harvard.edu
- [2505.21829] In Search of Adam's Secret Saucearxiv.org
- The Annotated Transformernlp.seas.harvard.edu
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Attention Is Off By One – Evan Millerevanmiller.org
- Online KL Shampoo | Tildeblog.tilderesearch.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Muon: An optimizer for hidden layers in neural networks | Keller Jordan blogkellerjordan.github.io
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io