TRAK-ing Model Behavior with Data – gradient science
gradientscience.org · 2,196 words · saved by 1 readers
Research highlights and perspectives on machine learning and optimization from MadryLab.
Homepage Code Paper In our latest paper, we revisit the problem of data attribution and introduce TRAK —a scalable and effective method for attributing machine learning predictions to training data. TRAK achieves dramatically better speed-efficacy tradeoffs than prior methods, allowing us to apply it across a variety of settings, from image classifiers to language models. As machine learning models become more capable (and simultaneously, more complex and opaque), the question of “why did my model make this prediction?” is becoming increasingly important. And the key to answering this question
related reading
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- I: Data problems (and solution concepts) in MLml-data-tutorial.org
- Attribution Patching: Activation Patching At Industrial Scale - Neel Nandaneelnanda.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- ADAG: Automatically Describing Attribution Graphsarxiv.org
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org
- Tracing Model Outputs to the Training Data \ Anthropicanthropic.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Bridging the Attention Gap: Complete Replacement Models for Complete Circuit Tracinginterp.open-moss.com
- Topicslearnmechinterp.com