Josh Engels on X: "New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! 🧵 https://t.co/mLg87XuXK5" / X
New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! 🧵 https://t.co/mLg87XuXK5
@JoshAEngels: New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! @JoshAEngels: This work is part of a series of research updates from our team: @JoshAEngels: My intuition for what’s going on: when we distill from a set of rollouts in SFT, we are establishing a strong prior on a certain assistant persona. Traits of that persona mostly aren’t explicitly optimized against i
Explore this link on the map →saved by
related reading
- SFT Drives Gemini’s Safety Properties — LessWronglesswrong.com
- SFT Drives Gemini’s Safety Properties — AI Alignment Forumalignmentforum.org
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- gpt-4.pdfcdn.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Expanding on what we missed with sycophancy | OpenAIopenai.com
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Rohin Shah on what it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | 80,000 Hours80000hours.org
- Thoughts on the impact of RLHF research — AI Alignment Forumalignmentforum.org
- Anthropic's leading researchers acted as moderate accelerationists — LessWronglesswrong.com