Anti-Slop Interventions? — LessWrong
In his recent post arguing against AI Control research, John Wentworth argues that the median doom path goes through AI slop, rather than scheming. I find this to be plausible. I believe this suggests a convergence of interests between AI capabilities research and AI alignment research. Historically, there has been a lot of concern about differential progress amongst AI safety researchers (perhaps especially those I tend to talk to). Some research is labeled as "capabilities" while other is labeled as "safety" (or, more often, "alignment"[1]). Most research is dual-use in practice (IE, has both capability and safety implications) and therefore should be kept secret or disclosed carefully. Recently, a colleague expressed concern that future AIs will read anything AI safety researchers publish now. Since the alignment of future AIs seems uncertain and even implausible, almost any information published now could be net harmful for the future. I argued the contrary case, as follows: a weak
x Anti-Slop Interventions? — LessWrong AI-Assisted Alignment Meta-Philosophy AI Frontpage 78 Anti-Slop Interventions? by abramdemski 4th Feb 2025 AI Alignment Forum 7 min read 33 78 Ω 26 In his recent post arguing against AI Control research, John Wentworth argues that the median doom path goes through AI slop, rather than scheming . I find this to be plausible. I believe this suggests a convergence of interests between AI capabilities research and AI alignment research. Historically, there has been a lot of concern about differential progress amongst AI safety researchers (perhaps especially
related reading
- The Case Against AI Control Research — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- The Universe from an Intentional Stancecasparoesterheld.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- The Problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Existential Risk from AI: An Exposition for Mathematiciansalkjash.github.io
- The Artificiality of Alignmentjoinreboot.org
- Reading Listblog.redwoodresearch.org