flâneur

SFT Drives Gemini’s Safety Properties — AI Alignment Forum

alignmentforum.org · 496 words · saved by 1 readers

This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adj…

x SFT Drives Gemini’s Safety Properties — AI Alignment Forum AI Frontpage 33 SFT Drives Gemini’s Safety Properties by Josh Engels , Arthur Conmy , bilalchughtai , Neel Nanda 13th Jun 2026 2 min read 4 33 This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here . In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages like RL. We do

related reading