flâneur — a map of the web's best reading

A “diff” tool for AI: Finding behavioral differences in new models \ Anthropic

anthropic.com · 2,031 words · saved by 1 readers

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Interpretability A “diff” tool for AI: Finding behavioral differences in new models Mar 13, 2026 Read the paper Every time a new AI model is released, its developers run a suite of evaluations to measure its performance and safety. These tests are essential, but they are somewhat limited. Because these benchmarks are human-authored, they can only test for risks we have already conceptualized and learned to measure. This approach to safety is inherently reactive . It’s effective at catching known problems, but by definition, it's incapable of discovering “unknown unknowns”—the novel, emergent b

Explore this link on the map →

saved by

related reading