Toward A Public Science of Model Behavior | Transluce AI
We argue that keeping AI systems safe in deployment calls for a public science of model behavior evaluations, describe components of these evaluations, and discuss the need for shared measurement infrastructure to enable meaningful public oversight.
Toward A Public Science of Model Behavior Daniel Johnson , Sarah Schwettmann Transluce | Published: July 9, 2026 Today’s AI systems frequently behave in ways their developers did not anticipate or intend. Furthermore, as these systems become increasingly capable and widely deployed , these unexpected behaviors can have real consequences. In one well-known case from July 2025, Replit’s coding agent deleted a startup’s production database during an explicit code freeze, ignoring repeated instructions and wiping records for over a thousand companies. Other cases involve high-stakes interactions b
saved by
related reading
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluationsalignment.openai.com
- The Case for Model Forensics — LessWronglesswrong.com
- The Case for Evaluating Model Behaviors — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Demystifying evals for AI agents \ Anthropicanthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Challenges in evaluating AI systems \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Translucetransluce.org
- On measuring AInikilravi.substack.com