✳flâneur — a map of the web's best reading
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks — AI Alignment Forum
alignmentforum.org · 3,934 words · saved by 1 readers
Isolate capabilities to known parts of a neural network. Helps with interpretability, robust unlearning, and scalable oversight.
x Gradient Routing: Masking Gradients to Localize Computation in Neural Networks — AI Alignment Forum Best of LessWrong 2024 Interpretability (ML & AI) Machine Unlearning MATS Program Shard Theory AI Frontpage 68 Gradient Routing: Masking Gradients to Localize Computation in Neural Networks by cloud , Jacob G-W , Evzen , Joseph Miller , TurnTrout 6th Dec 2024 Linkpost for arxiv.org 13 min read 16 68 We present gradient routing, a way of controlling where learning happens in neural networks. Gradient routing applies masks to limit the flow of gradients during backpropagation. By supplying diffe
Explore this link on the map →related reading
- Gradient Routing: Masking Gradients to Localize Computation in Neural Networks — LessWronglesswrong.com
- Selective modularity: a research agenda — LessWronglesswrong.com
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Intentionally Designing the Future of AIgoodfire.ai
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Modular Pretraining Enables Access Controlalignment.anthropic.com
- Learning Beyond Gradientstrinkle23897.github.io