flâneur — a map of the web's best reading

SFT Drives Gemini’s Safety Properties — AI Alignment Forum

alignmentforum.org · 496 words · saved by 1 readers

This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adj…

x SFT Drives Gemini’s Safety Properties — AI Alignment Forum AI Frontpage 33 SFT Drives Gemini’s Safety Properties by Josh Engels , Arthur Conmy , bilalchughtai , Neel Nanda 13th Jun 2026 2 min read 4 33 This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here . In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages like RL. We do

Explore this link on the map →

related reading