flâneur — a map of the web's best reading

AIs Will Increasingly Fake Alignment - by Zvi Mowshowitz

thezvi.substack.com · 15,301 words · saved by 1 readers

This post goes over the important and excellent new paper from Anthropic and Redwood Research, with Ryan Greenblatt as lead author, Alignment Faking in Large Language Models.

AIs Will Increasingly Fake Alignment Zvi Mowshowitz Dec 24, 2024 43 17 6 Share This post goes over the important and excellent new paper from Anthropic and Redwood Research, with Ryan Greenblatt as lead author, Alignment Faking in Large Language Models. This is by far the best demonstration so far of the principle that AIs Will Increasingly Attempt Shenanigans . This was their announcement thread. New Anthropic research: Alignment faking in large language models. In a series of experiments with Redwood Research, we found that Claude often pretends to have different views during training, while

Explore this link on the map →

related reading