The case for countermeasures to memetic spread of misaligned values
blog.redwoodresearch.org · 2,186 words · saved by 2 readers
Defending against alignment problems that might come with long-term memory
As various people have written about before, AIs that have long-term memory might pose additional risks (most notably, LLM AGI will have memory, and memory changes alignment by Seth Herd). Even if an AI is aligned or only occasionally scheming at the start of a deployment, the AI might become a consistent and coherent behavioral schemer via updates to its long-term memories. In this post, I’ll spell out the version of the threat model that I’m most concerned about, including some novel arguments for its plausibility, and describe some promising strategies for mitigating this risk. While I…
saved by
related reading
- The case for countermeasures to memetic spread of misaligned values — LessWronglesswrong.com
- When does training a model change its goals?blog.redwoodresearch.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Reading Listblog.redwoodresearch.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?blog.redwoodresearch.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?blog.redwoodresearch.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org