Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forum
This is the fourth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and ad…
x Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forum AI Frontpage 23 Why Do Naive SFT Filters For Safety Properties Fail? by Josh Engels , Neel Nanda 14th Jun 2026 13 min read 7 23 This is the fourth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The third post can be found here . Since SFT is the cause for many safety relevant properties , a natural strategy is to filter out rollouts from SFT that have undesirable properties. However, as we show in this section (and in forth
Explore this link on the map →saved by
related reading
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- SFT Drives Gemini’s Safety Properties — AI Alignment Forumalignmentforum.org
- SFT Drives Gemini’s Safety Properties — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Josh Engels on X: "New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! 🧵 https://t.co/mLg87XuXK5" / Xx.com
- gpt-4.pdfcdn.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com