Defining alignment research — LessWrong
I think that the concept of "alignment research" (and the distinction between that and "capabilities research") is currently a fairly confused one. In this post I’ll describe some of the problems with how people typically think about these terms, and offer replacement definitions. The first thing to highlight is that the distinction between alignment and capabilities is primarily doing useful work when we think of them as properties of AIs. This distinction is still under-appreciated by the wider machine learning community. ML researchers have historically thought about performance of models almost entirely with respect to the tasks they were specifically trained on. However, the rise of LLMs has vindicated the alignment community’s focus on general capabilities, and now it’s much more common to assume that performance on many tasks (including out-of-distribution tasks) will improve roughly in parallel. This is a crucial assumption for thinking about risks from AGI. Insofar as the ML c
x Defining alignment research — LessWrong AI Frontpage 131 Defining alignment research by Richard_Ngo 19th Aug 2024 AI Alignment Forum 9 min read 26 131 Ω 50 I think that the concept of "alignment research" (and the distinction between that and "capabilities research") is currently a fairly confused one. In this post I’ll describe some of the problems with how people typically think about these terms, and offer replacement definitions. “Alignment” and “capabilities” are primarily properties of AIs not of AI research The first thing to highlight is that the distinction between alignment and cap
saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Readings on the nature of alignment researchcasparoesterheld.com
- LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- The Universe from an Intentional Stancecasparoesterheld.com
- ALIGNMENT - by vincent huang - a slice of my mindmindslice.substack.com
- I Would Have Solved Alignment, But I Was Worried That Would Advance Timelines — LessWronglesswrong.com
- The Artificiality of Alignmentjoinreboot.org
- Tips for Empirical Alignment Research — AI Alignment Forumalignmentforum.org
- Differential acceleration of alignment-relevant capabilities is a bad bet — LessWronglesswrong.com
- A note about differential technological development — LessWronglesswrong.com
- Whose alignment research are we automating?firstscattering.com