✳flâneur — a map of the web's best reading
SFT Drives Gemini’s Safety Properties — AI Alignment Forum
alignmentforum.org · 496 words · saved by 1 readers
This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adj…
x SFT Drives Gemini’s Safety Properties — AI Alignment Forum AI Frontpage 33 SFT Drives Gemini’s Safety Properties by Josh Engels , Arthur Conmy , bilalchughtai , Neel Nanda 13th Jun 2026 2 min read 4 33 This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here . In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages like RL. We do
Explore this link on the map →related reading
- SFT Drives Gemini’s Safety Properties — LessWronglesswrong.com
- Josh Engels on X: "New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! 🧵 https://t.co/mLg87XuXK5" / Xx.com
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- gpt-4.pdfcdn.openai.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Rohin Shah on what it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | 80,000 Hours80000hours.org
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- ARENA - AI Safety Curriculumlearn.arena.education
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org