flâneur — a map of the web's best reading

Anti-Slop Interventions? — LessWrong

lesswrong.com · 10,073 words · saved by 1 readers

In his recent post arguing against AI Control research, John Wentworth argues that the median doom path goes through AI slop, rather than scheming. I find this to be plausible. I believe this suggests a convergence of interests between AI capabilities research and AI alignment research. Historically, there has been a lot of concern about differential progress amongst AI safety researchers (perhaps especially those I tend to talk to). Some research is labeled as "capabilities" while other is labeled as "safety" (or, more often, "alignment"[1]). Most research is dual-use in practice (IE, has both capability and safety implications) and therefore should be kept secret or disclosed carefully. Recently, a colleague expressed concern that future AIs will read anything AI safety researchers publish now. Since the alignment of future AIs seems uncertain and even implausible, almost any information published now could be net harmful for the future. I argued the contrary case, as follows: a weak

x Anti-Slop Interventions? — LessWrong AI-Assisted Alignment Meta-Philosophy AI Frontpage 78 Anti-Slop Interventions? by abramdemski 4th Feb 2025 AI Alignment Forum 7 min read 33 78 Ω 26 In his recent post arguing against AI Control research, John Wentworth argues that the median doom path goes through AI slop, rather than scheming . I find this to be plausible. I believe this suggests a convergence of interests between AI capabilities research and AI alignment research. Historically, there has been a lot of concern about differential progress amongst AI safety researchers (perhaps especially

Explore this link on the map →

related reading