✳flâneur — a map of the web's best reading
The case for countermeasures to memetic spread of misaligned values — LessWrong
lesswrong.com · 4,428 words · saved by 1 readers
As various people have written about before, AIs that have long-term memory might pose additional risks (most notably, LLM AGI will have memory, and…
x The case for countermeasures to memetic spread of misaligned values — LessWrong Redwood Research AI Frontpage 83 The case for countermeasures to memetic spread of misaligned values by Alex Mallen 28th May 2025 AI Alignment Forum 8 min read 8 83 Ω 46 As various people have written about before, AIs that have long-term memory might pose additional risks (most notably, LLM AGI will have memory, and memory changes alignment by Seth Herd). Even if an AI is aligned or only occasionally scheming at the start of a deployment, the AI might become a consistent and coherent behavioral schemer via updat
Explore this link on the map →saved by
related reading
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- We're already in AI takeoff — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Spring 2026 Projects - SPARsparai.org
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- Why You Don’t Believe in Xhosa Prophecies — LessWronglesswrong.com
- Training-time schemers vs behavioral schemers — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org