flâneur — a map of the web's best reading

Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forum

alignmentforum.org · 3,230 words · saved by 1 readers

This is the fourth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and ad…

x Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forum AI Frontpage 23 Why Do Naive SFT Filters For Safety Properties Fail? by Josh Engels , Neel Nanda 14th Jun 2026 13 min read 7 23 This is the fourth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The third post can be found here . Since SFT is the cause for many safety relevant properties , a natural strategy is to filter out rollouts from SFT that have undesirable properties. However, as we show in this section (and in forth

Explore this link on the map →

saved by

related reading