Toward A Public Science of Model Behavior | Transluce AI
We argue that keeping AI systems safe in deployment calls for a public science of model behavior evaluations, describe components of these evaluations, and discuss the need for shared measurement infrastructure to enable meaningful public oversight.
Toward A Public Science of Model Behavior Daniel Johnson , Sarah Schwettmann Transluce | Published: July 9, 2026 Today’s AI systems frequently behave in ways their developers did not anticipate or intend. Furthermore, as these systems become increasingly capable and widely deployed , these unexpected behaviors can have real consequences. In one well-known case from July 2025, Replit’s coding agent deleted a startup’s production database during an explicit code freeze, ignoring repeated instructions and wiping records for over a thousand companies. Other cases involve high-stakes interactions b
Explore this link on the map →related reading
- The Case for Evaluating Model Behaviors — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- We need a Science of Evals — AI Alignment Forumalignmentforum.org
- We Need A ‘Science of Evals’ – Apollo Researchapolloresearch.ai
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- AI Safety Seems Hard to Measurecold-takes.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluationsarxiv.org