flâneur — a map of the web's best reading

Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment — LessWrong

lesswrong.com · 2,833 words · saved by 1 readers

Background Deliberative alignment is a powerful post-training alignment technique that involves generating and training on re-contextualised supervis…

x Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment — LessWrong AI Frontpage 54 Generalisation Hacking: a first look at adversarial generalisation failures in deliberative alignment by Cam , Puria 17th Nov 2025 AI Alignment Forum 10 min read 2 54 Ω 16 Background Deliberative alignment is a powerful post-training alignment technique that involves generating and training on re-contextualised supervised fine-tuning (SFT) datasets generated with a set of principles in context. The process takes three steps: With the set of principles [1] (hencefo

Explore this link on the map →

related reading