✳flâneur — a map of the web's best reading
SFT Drives Gemini’s Safety Properties — LessWrong
lesswrong.com · 807 words · saved by 1 readers
This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adj…
x SFT Drives Gemini’s Safety Properties — LessWrong AI Frontpage 89 SFT Drives Gemini’s Safety Properties by Josh Engels , Arthur Conmy , bilalchughtai , Neel Nanda 13th Jun 2026 AI Alignment Forum 2 min read 4 89 Ω 33 This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here . In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages
Explore this link on the map →related reading
- SFT Drives Gemini’s Safety Properties — AI Alignment Forumalignmentforum.org
- Josh Engels on X: "New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! 🧵 https://t.co/mLg87XuXK5" / Xx.com
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- gpt-4.pdfcdn.openai.com
- AI in 2025: gestalt — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- ARENA - AI Safety Curriculumlearn.arena.education
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Synthetic Persona Pretraining: Alignment from Token Zero — LessWronglesswrong.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com