Influence functions - why, what and how — LessWrong
Anthropic recently published the paper Studying Large Language Model Generalization with Influence Functions, which describes a scalable technique fo…
x Influence functions - why, what and how — LessWrong Machine Learning (ML) AI Frontpage 78 Influence functions - why, what and how by Nina Panickssery 15th Sep 2023 10 min read 6 78 Anthropic recently published the paper Studying Large Language Model Generalization with Influence Functions , which describes a scalable technique for measuring which training examples were most influential for a particular set of weights/outputs of a trained model. This can help us better understand model generalization, offering insights into the emergent properties of AI systems. For instance, influence functi
Explore this link on the map →related reading
- [2308.03296] Studying Large Language Model Generalization with Influence Functionsarxiv.org
- Tracing Model Outputs to the Training Data \ Anthropicanthropic.com
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- The Little Book of Deep Learningfleuret.org
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- microgptkarpathy.github.io
- Just Ask for Generalization | Eric Jangevjang.com
- arxiv.org/pdf/1805.08522arxiv.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- [2606.04071] Covert Influence Between Language Modelsarxiv.org