Caspar Oesterheld
casparoesterheld.com · 9,379 words · saved by 2 readers
A blog on Philosophy, Artificial Intelligence and Effective Altruism
Overview: Relative to AI capabilities research, AI alignment research seems more conceptual/fuzzy/… I’ll first try to express a version of this point myself. But primarily this post serves as a mildly curated reading list of texts by others making similar points. I’ll also give some examples of conceptual/fuzzy/… work in alignment research. In my mind, an important fact about alignment research (e.g., for figuring out how to automate it, for how to design work tests, how to organize the research community) is that it has in part the following cluster of properties: preparadigmatic;…
saved by
related reading
- LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Shtetl-Optimized >> Blog Archive >> Theory and AI Alignmentscottaaronson.blog
- Readings on the nature of alignment researchcasparoesterheld.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- I Would Have Solved Alignment, But I Was Worried That Would Advance Timelines — LessWronglesswrong.com
- Defining alignment research — LessWronglesswrong.com
- The Artificiality of Alignmentjoinreboot.org
- The Best of LessWrong — LessWronglesswrong.com
- Differential acceleration of alignment-relevant capabilities is a bad bet — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com