flâneur

The case for countermeasures to memetic spread of misaligned values

blog.redwoodresearch.org · 2,186 words · saved by 2 readers

Defending against alignment problems that might come with long-term memory

As various people have written about before, AIs that have long-term memory might pose additional risks (most notably, LLM AGI will have memory, and memory changes alignment by Seth Herd). Even if an AI is aligned or only occasionally scheming at the start of a deployment, the AI might become a consistent and coherent behavioral schemer via updates to its long-term memories. In this post, I’ll spell out the version of the threat model that I’m most concerned about, including some novel arguments for its plausibility, and describe some promising strategies for mitigating this risk. While I…

saved by

related reading