✳flâneur — a map of the web's best reading
A “diff” tool for AI: Finding behavioral differences in new models \ Anthropic
anthropic.com · 2,031 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Interpretability A “diff” tool for AI: Finding behavioral differences in new models Mar 13, 2026 Read the paper Every time a new AI model is released, its developers run a suite of evaluations to measure its performance and safety. These tests are essential, but they are somewhat limited. Because these benchmarks are human-authored, they can only test for risks we have already conceptualized and learned to measure. This approach to safety is inherently reactive . It’s effective at catching known problems, but by definition, it's incapable of discovering “unknown unknowns”—the novel, emergent b
Explore this link on the map →saved by
related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Insights on Crosscoder Model Diffingtransformer-circuits.pub
- How confessions can keep language models honest | OpenAIopenai.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- What We Learned Trying to Diff Base and Chat Models (And Why It Matters) — LessWronglesswrong.com
- Stage-Wise Model Diffingtransformer-circuits.pub
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Claude 4 System Cardwww-cdn.anthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- [2602.22755] AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviorsarxiv.org
- AI #77: A Few Upgrades - by Zvi Mowshowitzthezvi.substack.com