flâneur — a map of the web's best reading

How To Go From Interpretability To Alignment: Just Retarget The Search — LessWrong

lesswrong.com · 6,650 words · saved by 2 readers

Here's a simple strategy for AI alignment: use interpretability tools to identify the AI's internal search process, and the AI's internal representat…

x How To Go From Interpretability To Alignment: Just Retarget The Search — LessWrong "Why Not Just..." Inner Alignment Interpretability (ML & AI) AI Risk AI Frontpage 214 How To Go From Interpretability To Alignment: Just Retarget The Search by johnswentworth 10th Aug 2022 AI Alignment Forum 3 min read 34 214 Ω 76 [EDIT: Many people who read this post were very confused about some things, which I later explained in What’s General-Purpose Search, And Why Might We Expect To See It In Trained ML Systems? You might want to read that post first.] When people talk about prosaic alignment proposals ,

Explore this link on the map →

saved by

related reading