flâneur — a map of the web's best reading

Defining alignment research — LessWrong

lesswrong.com · 5,899 words · saved by 1 readers

I think that the concept of "alignment research" (and the distinction between that and "capabilities research") is currently a fairly confused one. In this post I’ll describe some of the problems with how people typically think about these terms, and offer replacement definitions. The first thing to highlight is that the distinction between alignment and capabilities is primarily doing useful work when we think of them as properties of AIs. This distinction is still under-appreciated by the wider machine learning community. ML researchers have historically thought about performance of models almost entirely with respect to the tasks they were specifically trained on. However, the rise of LLMs has vindicated the alignment community’s focus on general capabilities, and now it’s much more common to assume that performance on many tasks (including out-of-distribution tasks) will improve roughly in parallel. This is a crucial assumption for thinking about risks from AGI. Insofar as the ML c

x Defining alignment research — LessWrong AI Frontpage 131 Defining alignment research by Richard_Ngo 19th Aug 2024 AI Alignment Forum 9 min read 26 131 Ω 50 I think that the concept of "alignment research" (and the distinction between that and "capabilities research") is currently a fairly confused one. In this post I’ll describe some of the problems with how people typically think about these terms, and offer replacement definitions. “Alignment” and “capabilities” are primarily properties of AIs not of AI research The first thing to highlight is that the distinction between alignment and cap

Explore this link on the map →

related reading