SFT Drives Gemini’s Safety Properties — LessWrong
lesswrong.com · 807 words · saved by 5 readers
This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adj…
x SFT Drives Gemini’s Safety Properties — LessWrong AI Frontpage 89 SFT Drives Gemini’s Safety Properties by Josh Engels , Arthur Conmy , bilalchughtai , Neel Nanda 13th Jun 2026 AI Alignment Forum 2 min read 4 89 Ω 33 This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here . In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages
saved by
related reading
- Josh Engels on X: "New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! 🧵 https://t.co/mLg87XuXK5" / Xx.com
- SFT Drives Gemini’s Safety Properties — AI Alignment Forumalignmentforum.org
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- When Should We Introduce Safety Interventions During Pretraining?arxiv.org
- ARENA - AI Safety Curriculumlearn.arena.education
- Dylan Samdsam99.github.io
- GPT-6 Astra System Card - OpenAI Deployment Safety Hubdeploymentsafety.openai.com
- gpt-4.pdfcdn.openai.com
- Papers and Projects - Josh Engelsjoshengels.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- [2510.27062] Consistency Training Helps Stop Sycophancy and Jailbreaksarxiv.org