Selective modularity: a research agenda — LessWrong
lesswrong.com · 7,853 words · saved by 1 readers
By training neural networks with selective modularity, gradient routing enables new approaches to core problems in AI safety.
x Selective modularity: a research agenda — LessWrong AI Frontpage 72 Selective modularity: a research agenda by cloud , Jacob G-W 24th Mar 2025 AI Alignment Forum 29 min read 3 72 Ω 29 Overview: By training neural networks with selective modularity, gradient routing enables new approaches to core problems in AI safety. This agenda identifies related research directions that might enable safer development of transformative AI. Introduction Soon, the world may see rapid increases in AI capabilities resulting from AI research automation, and no one knows how to ensure this happens safely ( Soare
related reading
- 2410.04332arxiv.org
- Gradient Routing: Masking Gradients to Localize Computation in Neural Networks — AI Alignment Forumalignmentforum.org
- Gradient Routing: Masking Gradients to Localize Computation in Neural Networks — LessWronglesswrong.com
- Modular Pretraining Enables Access Controlalignment.anthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Bitter Lessons from Distillation Robustifies Unlearningbrucewlee.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Intentionally Designing the Future of AIgoodfire.ai
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- In (highly contingent!) defense of interpretability-in-the-loop ML training — AI Alignment Forumalignmentforum.org