Anti-Slop Interventions? — LessWrong
In his recent post arguing against AI Control research, John Wentworth argues that the median doom path goes through AI slop, rather than scheming. I find this to be plausible. I believe this suggests a convergence of interests between AI capabilities research and AI alignment research. Historically, there has been a lot of concern about differential progress amongst AI safety researchers (perhaps especially those I tend to talk to). Some research is labeled as "capabilities" while other is labeled as "safety" (or, more often, "alignment"[1]). Most research is dual-use in practice (IE, has both capability and safety implications) and therefore should be kept secret or disclosed carefully. Recently, a colleague expressed concern that future AIs will read anything AI safety researchers publish now. Since the alignment of future AIs seems uncertain and even implausible, almost any information published now could be net harmful for the future. I argued the contrary case, as follows: a weak
x Anti-Slop Interventions? — LessWrong AI-Assisted Alignment Meta-Philosophy AI Frontpage 78 Anti-Slop Interventions? by abramdemski 4th Feb 2025 AI Alignment Forum 7 min read 33 78 Ω 26 In his recent post arguing against AI Control research, John Wentworth argues that the median doom path goes through AI slop, rather than scheming . I find this to be plausible. I believe this suggests a convergence of interests between AI capabilities research and AI alignment research. Historically, there has been a lot of concern about differential progress amongst AI safety researchers (perhaps especially
Explore this link on the map →related reading
- The Case Against AI Control Research — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- The Best of LessWrong — LessWronglesswrong.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- I Would Have Solved Alignment, But I Was Worried That Would Advance Timelines — LessWronglesswrong.com
- AI Pause Will Likely Backfire — EA Forumforum.effectivealtruism.org