✳flâneur — a map of the web's best reading
Adam Optimizer Causes Privileged Basis in Transformer Language Models — LessWrong
lesswrong.com · 1,535 words · saved by 1 readers
Diego Caples (diego@activated-ai.com) • Rob Neuhaus (rob@activated-ai.com) …
x Adam Optimizer Causes Privileged Basis in Transformer LM Residual Stream — LessWrong Interpretability (ML & AI) Machine Learning (ML) Optimization Transformers AI Frontpage 74 Adam Optimizer Causes Privileged Basis in Transformer LM Residual Stream by Diego Caples , rrenaud 6th Sep 2024 4 min read 8 74 Diego Caples ( diego@activated-ai.com ) Rob Neuhaus ( rob@activated-ai.com ) Introduction In principle, neuron activations in a transformer-based language model residual stream should be about the same scale. In practice, the dimensions unexpectedly widely vary in scale. Mathematical theories
Explore this link on the map →related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- [2505.21829] In Search of Adam's Secret Saucearxiv.org
- The Annotated Transformernlp.seas.harvard.edu
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai
- Attention Is Off By One – Evan Millerevanmiller.org
- The Annotated Transformernlp.seas.harvard.edu
- Online KL Shampoo | Tildeblog.tilderesearch.com
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io