flâneur — a map of the web's best reading

Tracing Model Outputs to the Training Data \ Anthropic

anthropic.com · 1,089 words · saved by 1 readers

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Alignment Research Tracing Model Outputs to the Training Data Aug 8, 2023 As large language models become more powerful and their risks become clearer, there is increasing value to figuring out what makes them tick. In our previous work , we have found that large language models change along many personality and behavioral dimensions as a function of both scale and the amount of fine-tuning. Understanding these changes requires seeing how models work, for instance to determine if a model’s outputs rely on memorization or more sophisticated processing. Understanding the inner workings of langua

Explore this link on the map →

related reading