The case for countermeasures to memetic spread of misaligned values — LessWrong
lesswrong.com · 4,428 words · saved by 1 readers
As various people have written about before, AIs that have long-term memory might pose additional risks (most notably, LLM AGI will have memory, and…
x The case for countermeasures to memetic spread of misaligned values — LessWrong Redwood Research AI Frontpage 83 The case for countermeasures to memetic spread of misaligned values by Alex Mallen 28th May 2025 AI Alignment Forum 8 min read 8 83 Ω 46 As various people have written about before, AIs that have long-term memory might pose additional risks (most notably, LLM AGI will have memory, and memory changes alignment by Seth Herd). Even if an AI is aligned or only occasionally scheming at the start of a deployment, the AI might become a consistent and coherent behavioral schemer via updat
saved by
related reading
- The case for countermeasures to memetic spread of misaligned valuesblog.redwoodresearch.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- We're already in AI takeoff — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?blog.redwoodresearch.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?blog.redwoodresearch.org