✳flâneur — a map of the web's best reading
AI Interpretability — ML Alignment & Theory Scholars
matsprogram.org · saved by 1 readers
Rigorously understanding how ML models function may allow us to identify and train against misalignment. Can we reverse engineer neural nets from their weights, or identify structures corresponding to “goals” or dangerous capabilities within a model and surgically alter them?
Explore this link on the map →