Can public chat data predict real-world AI misalignments?
alignment.openai.com · 2,955 words · saved by 1 readers
Bridging private deployment evidence and public AI evaluation
Frontier AI models are increasingly used in settings with real economic, legal, and societal consequences. As a result, governments, AI safety organizations and independent researchers need ways to evaluate how these systems behave under realistic conditions. Traditional evaluations use hand-written, synthetic, or adversarial prompts to stress-test known risks and compare models under controlled conditions. But these prompts can be narrow, unrepresentative, or recognizable as tests. An alternative, complementary way to evaluate how models behave in the real world is often to look at real…
saved by
related reading
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- Teaching Claude Whyalignment.anthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Today, we are releasing a research preview of our user model, along with a set of evaluations designed to measure how faithfully user models capture human behavior.persimmon.humansand.ai
- Predicting model behavior before release by simulating deployment | OpenAIopenai.com
- Toward A Public Science of Model Behavior | Transluce AItransluce.org
- Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluationsalignment.openai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- 2405.01470arxiv.org
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Frontier Risk Report (February to March 2026) - METRmetr.org