flâneur — a map of the web's best reading

SFT Drives Gemini’s Safety Properties — LessWrong

lesswrong.com · 807 words · saved by 1 readers

This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adj…

x SFT Drives Gemini’s Safety Properties — LessWrong AI Frontpage 89 SFT Drives Gemini’s Safety Properties by Josh Engels , Arthur Conmy , bilalchughtai , Neel Nanda 13th Jun 2026 AI Alignment Forum 2 min read 4 89 Ω 33 This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here . In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages

Explore this link on the map →

related reading