Explaining AI Alignment as an NLPer and Why I am Working on It
I started NLP in 2018 and worked on various NLP topics related to algorithmic bias, interpretability, and semantic parsing. In 2021, I pivoted to Scalable Oversight, a sub-area of AI alignment.
I started NLP in 2018 and worked on various NLP topics related to algorithmic bias, interpretability, and semantic parsing. In 2021, I pivoted to Scalable Oversight, a sub-area of AI alignment. Since the successes of chatbots such as GPT-4 and Claude, more NLP researchers start to discuss AI alignment. However, people use the word “AI alignment” for different things: some think it means “preventing AI systems from taking over the world” while others think it specifically means “optimizing for human ratings”. To clarify these misconceptions, I will first introduce one particular definition…
saved by
related reading
- Automated alignment is harder than you thinkarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- The Artificiality of Alignmentjoinreboot.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Defining alignment research — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com