flâneur — a map of the web's best reading

AI Interpretability — ML Alignment & Theory Scholars

matsprogram.org · saved by 1 readers

Rigorously understanding how ML models function may allow us to identify and train against misalignment. Can we reverse engineer neural nets from their weights, or identify structures corresponding to “goals” or dangerous capabilities within a model and surgically alter them?

Explore this link on the map →

saved by