flâneur

Can public chat data predict real-world AI misalignments?

alignment.openai.com · 2,955 words · saved by 1 readers

Bridging private deployment evidence and public AI evaluation

Frontier AI models are increasingly used in settings with real economic, legal, and societal consequences. As a result, governments, AI safety organizations and independent researchers need ways to evaluate how these systems behave under realistic conditions. Traditional evaluations use hand-written, synthetic, or adversarial prompts to stress-test known risks and compare models under controlled conditions. But these prompts can be narrow, unrepresentative, or recognizable as tests. An alternative, complementary way to evaluate how models behave in the real world is often to look at real…

saved by

related reading