✳flâneur — a map of the web's best reading
Tracing Model Outputs to the Training Data \ Anthropic
anthropic.com · 1,089 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Alignment Research Tracing Model Outputs to the Training Data Aug 8, 2023 As large language models become more powerful and their risks become clearer, there is increasing value to figuring out what makes them tick. In our previous work , we have found that large language models change along many personality and behavioral dimensions as a function of both scale and the amount of fine-tuning. Understanding these changes requires seeing how models work, for instance to determine if a model’s outputs rely on memorization or more sophisticated processing. Understanding the inner workings of langua
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- On the Biology of a Large Language Modeltransformer-circuits.pub
- [2308.03296] Studying Large Language Model Generalization with Influence Functionsarxiv.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Tracing the Thoughts of a Large Language Model — LessWronglesswrong.com
- Influence functions - why, what and how — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- [2606.04071] Covert Influence Between Language Modelsarxiv.org
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com