✳flâneur — a map of the web's best reading
Selective modularity: a research agenda — LessWrong
lesswrong.com · 7,853 words · saved by 1 readers
By training neural networks with selective modularity, gradient routing enables new approaches to core problems in AI safety.
x Selective modularity: a research agenda — LessWrong AI Frontpage 72 Selective modularity: a research agenda by cloud , Jacob G-W 24th Mar 2025 AI Alignment Forum 29 min read 3 72 Ω 29 Overview: By training neural networks with selective modularity, gradient routing enables new approaches to core problems in AI safety. This agenda identifies related research directions that might enable safer development of transformative AI. Introduction Soon, the world may see rapid increases in AI capabilities resulting from AI research automation, and no one knows how to ensure this happens safely ( Soare
Explore this link on the map →related reading
- Gradient Routing: Masking Gradients to Localize Computation in Neural Networks — AI Alignment Forumalignmentforum.org
- Gradient Routing: Masking Gradients to Localize Computation in Neural Networks — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- Intentionally Designing the Future of AIgoodfire.ai
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Modular Pretraining Enables Access Controlalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com