Narrow finetuning is different — LessWrong
lesswrong.com · 1,468 words · saved by 3 readers
It is common to use finetuning on a narrow data distribution, or narrow finetuning (NFT), to study AI safety. In these experiments, a model is traine…
x Narrow finetuning is different — LessWrong AI Frontpage 70 Narrow finetuning is different by cloud , Stewy Slocum 5th Aug 2025 4 min read 3 70 It is common to use finetuning on a narrow data distribution, or narrow finetuning (NFT), to study AI safety. In these experiments, a model is trained on a very specific type of data, then evaluated for broader properties, such as a capability or general disposition. Ways that narrow finetuning is different Narrow finetuning is different than the training procedures that frontier AI companies use, like pretraining on the internet , or posttraining on
saved by
related reading
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- Conditioning, Prompts, and Fine-Tuning — LessWronglesswrong.com
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- [2606.00831] Subliminal Learning is a LoRA Artifactarxiv.org
- Anatomy of a Modern Finetuning APIbenanderson.work
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org