✳flâneur — a map of the web's best reading
Narrow finetuning is different — LessWrong
lesswrong.com · 1,468 words · saved by 2 readers
It is common to use finetuning on a narrow data distribution, or narrow finetuning (NFT), to study AI safety. In these experiments, a model is traine…
x Narrow finetuning is different — LessWrong AI Frontpage 70 Narrow finetuning is different by cloud , Stewy Slocum 5th Aug 2025 4 min read 3 70 It is common to use finetuning on a narrow data distribution, or narrow finetuning (NFT), to study AI safety. In these experiments, a model is trained on a very specific type of data, then evaluated for broader properties, such as a capability or general disposition. Ways that narrow finetuning is different Narrow finetuning is different than the training procedures that frontier AI companies use, like pretraining on the internet , or posttraining on
Explore this link on the map →saved by
related reading
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Generalization Dynamics of LM Pre-training — Jiaxin Wenjiaxin-wen.github.io
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Anatomy of a Modern Finetuning APIbenanderson.work
- How far does alignment midtraining generalize?alignment.openai.com
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs 49 This paper contains model-generated content that might be offensive. 49arxiv.org
- Modern Pretraining Strategies: A Hands-On Guidetheneuralmaze.substack.com
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences — LessWronglesswrong.com