SFT Drives Gemini’s Safety Properties — AI Alignment Forum
alignmentforum.org · 496 words · saved by 1 readers
This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adj…
x SFT Drives Gemini’s Safety Properties — AI Alignment Forum AI Frontpage 33 SFT Drives Gemini’s Safety Properties by Josh Engels , Arthur Conmy , bilalchughtai , Neel Nanda 13th Jun 2026 2 min read 4 33 This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here . In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages like RL. We do
related reading
- SFT Drives Gemini’s Safety Properties — LessWronglesswrong.com
- Josh Engels on X: "New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! 🧵 https://t.co/mLg87XuXK5" / Xx.com
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- gpt-4.pdfcdn.openai.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- GPT-6 Astra System Carddeploymentsafety.openai.com
- A Summary of Recent Work (July 2026)gdmalignment.substack.com
- Frontier Safety Framework Report - Gemini 3 Pro (November, 2025) v2storage.googleapis.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org