flâneur — a map of the web's best reading

“Alignment Faking” frame is somewhat fake — LessWrong

lesswrong.com · 5,531 words · saved by 1 readers

I like the research. I mostly trust the results. I dislike the 'Alignment Faking' name and frame, and I'm afraid it will stick and lead to more confu…

x “Alignment Faking” frame is somewhat fake — LessWrong Best of LessWrong 2024 Deceptive Alignment AI Frontpage 170 “Alignment Faking” frame is somewhat fake by Jan_Kulveit 20th Dec 2024 AI Alignment Forum 7 min read 16 170 Ω 73 I like the research. I mostly trust the results. I dislike the 'Alignment Faking' name and frame, and I'm afraid it will stick and lead to more confusion. This post offers a different frame. The main way I think about the result is: it's about capability - the model exhibits strategic preference preservation behavior ; also, harmlessness generalized better than honesty

Explore this link on the map →

related reading