2410.04332
arxiv.org · 7,905 words · saved by 1 readers
N/A
Preprint G RADIENT ROUTING : M ASKING G RADIENTS TO L O - CALIZE C OMPUTATION IN N EURAL N ETWORKS Alex Cloud∗ , Jacob Goldman-Wetzler∗ , Evžen Wybitul∗ , Joseph Miller∗ ML Alignment and Theory Scholars (MATS) Alexander Matt Turner arXiv:2410.04332v2 [cs.LG] 29 Nov 2024 A BSTRACT…
related reading
- Gradient Routing: Masking Gradients to Localize Computation in Neural Networks — AI Alignment Forumalignmentforum.org
- Gradient Routing: Masking Gradients to Localize Computation in Neural Networks — LessWronglesswrong.com
- Selective modularity: a research agenda — LessWronglesswrong.com
- Modular Pretraining Enables Access Controlalignment.anthropic.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- [2512.05648] Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMsarxiv.org
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forumalignmentforum.org
- Interpreting Language Model Parametersgoodfire.ai
- A Comprehensive Mechanistic Interpretability Explainer & Glossary — Neel Nandaneelnanda.io