[2308.03296] Studying Large Language Model Generalization with Influence Functions
Abstract:When trying to gain better visibility into a machine learning model in order to understand and mitigate the associated risks, a potentially valuable source of evidence is: which training examples most contribute to a given behavior? Influence functions aim to answer a counterfactual: how would the model's parameters (and hence its outputs) change if a given sequence were added to the training set? While influence functions have produced insights for small models, they are difficult to scale to large language models (LLMs) due to the difficulty of computing an inverse-Hessian-vector product (IHVP). We use the Eigenvalue-corrected Kronecker-Factored Approximate Curvature (EK-FAC) approximation to scale influence functions up to LLMs with up to 52 billion parameters. In our experiments, EK-FAC achieves similar accuracy to traditional influence function estimators despite the IHVP computation being orders of magnitude faster. We investigate two algorithmic techniques to reduce the cost of computing gradients of candidate training sequences: TF-IDF filtering and query batching. We use influence functions to investigate the generalization patterns of LLMs, including the sparsity of the influence patterns, increasing abstraction with scale, math and programming abilities, cross-lingual generalization, and role-playing behavior. Despite many apparently sophisticated forms of generalization, we identify a surprising limitation: influences decay to near-zero when the order of key phrases is flipped. Overall, influence functions give us a powerful new tool for studying the generalization properties of LLMs.
[2308.03296] Studying Large Language Model Generalization with Influence Functions Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:2308.03296 (cs) [Submitted on 7 Aug 2023] Title: Studying Large Language Model Generalization with Influence Functions Authors: Roger Grosse , Juhan Bae , Cem Anil , Nelson Elhage , Alex Tamkin , Amirhossein Tajdini , Benoit Steiner , Dustin Li , Esin Durmus , Ethan Perez , Evan Hubinger , Kamilė Lukošiūtė , Karina Nguyen , Nichol
Explore this link on the map →related reading
- Influence functions - why, what and how — LessWronglesswrong.com
- Tracing Model Outputs to the Training Data \ Anthropicanthropic.com
- Just Ask for Generalization | Eric Jangevjang.com
- GenAI Handbookgenai-handbook.github.io
- Large Language Model: world models or surface statistics?thegradient.pub
- LLM Resourcesforrestbicker.com
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Are Emergent Abilities in Large Language Models just In-Context Learning?arxiv.org
- Anthropic on X: "Here is another example of increasing abstraction with scale, where an AI Assistant reasoned through an AI alignment question. The top influential sequence for the 810M model shares a short phrase with the query, while thetwitter.com
- arxiv.org/pdf/2505.24832arxiv.org