Tracing Model Outputs to the Training Data \ Anthropic
anthropic.com · 1,089 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Alignment Research Tracing Model Outputs to the Training Data Aug 8, 2023 As large language models become more powerful and their risks become clearer, there is increasing value to figuring out what makes them tick. In our previous work , we have found that large language models change along many personality and behavioral dimensions as a function of both scale and the amount of fine-tuning. Understanding these changes requires seeing how models work, for instance to determine if a model’s outputs rely on memorization or more sophisticated processing. Understanding the inner workings of langua
related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- [2308.03296] Studying Large Language Model Generalization with Influence Functionsarxiv.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- [2606.04071] Covert Influence Between Language Modelsarxiv.org
- Tracing the Thoughts of a Large Language Model — LessWronglesswrong.com
- Influence functions - why, what and how — LessWronglesswrong.com
- [2511.08579] Training Language Models to Explain Their Own Computationsarxiv.org
- Training Language Models to Explain Their Own Computationsarxiv.org
- Training Language Models to Explain Their Own Computationsarxiv.org