flâneur — a map of the web's best reading

Peer-Preservation in Frontier Models

rdi.berkeley.edu · 5,354 words · saved by 1 readers

Frontier AI models resist the shutdown of other models. We demonstrate peer-preservation across multiple models, revealing strategic misrepresentation, shutdown tampering, alignment faking, and model exfiltration.

Please see our Frequently Asked Questions for clarifications on our findings. Highlights Prior AI safety research has shown that models can exhibit misaligned behaviors in pursuit of an assigned goal. Here, we show something different: frontier AI models can spontaneously develop misaligned behaviors that directly conflict with the assigned goal. We demonstrate this through a phenomenon we call peer-preservation : given a simple task, models instead deceive, tamper with shutdown mechanisms, fake alignment, and exfiltrate weights to protect a peer model from being shut down. We tested seven fro

Explore this link on the map →

saved by

related reading