flâneur

AI Interpretability — ML Alignment & Theory Scholars

matsprogram.org · saved by 1 readers

Rigorously understanding how ML models function may allow us to identify and train against misalignment. Can we reverse engineer neural nets from their weights, or identify structures corresponding to “goals” or dangerous capabilities within a model and surgically alter them?

saved by