flâneur — a map of the web's best reading

Investigating truthfulness in a pre-release o3 model | Transluce AI

transluce.org · 5,240 words · saved by 1 readers

During pre-release testing of OpenAI's o3 model, we found that o3 frequently fabricates actions it took to fulfill user requests, and elaborately justifies the fabrications when confronted by the user. These behaviors generalize to other reasoning models that are widely deployed, such as o3-mini and o1. To dig deeper, we automatically generated hundreds of example behaviors and analyzed them with Docent, surfacing other unexpected behaviors including the model's disclosure of the "Yap score" in its system message. Transluce received early testing access to OpenAI's frontier o3 model. We tested the model (o3-2025-04-03) with a combination of human users and investigator agents that learned from the human users, to scale the number and types of possible interactions. We also used Docent to analyze the resulting transcripts for surprising behaviors. We surfaced a variety of behaviors that indicate truthfulness issues in o3. Most frequently, o3 hallucinates usage of a code tool and associa

Investigating truthfulness in a pre-release o3 model Neil Chowdhury* , Daniel Johnson , Vincent Huang , Jacob Steinhardt , Sarah Schwettmann* * Correspondence to: neil@transluce.org, sarah@transluce.org Transluce | Published: April 16, 2025 During pre-release testing of OpenAI's o3 model, we found that o3 frequently fabricates actions it took to fulfill user requests, and elaborately justifies the fabrications when confronted by the user. These behaviors generalize to other reasoning models that are widely deployed, such as o3-mini and o1. To dig deeper, we automatically generated hundreds of

Explore this link on the map →

related reading