Investigating truthfulness in a pre-release o3 model | Transluce AI
During pre-release testing of OpenAI's o3 model, we found that o3 frequently fabricates actions it took to fulfill user requests, and elaborately justifies the fabrications when confronted by the user. These behaviors generalize to other reasoning models that are widely deployed, such as o3-mini and o1. To dig deeper, we automatically generated hundreds of example behaviors and analyzed them with Docent, surfacing other unexpected behaviors including the model's disclosure of the "Yap score" in its system message. Transluce received early testing access to OpenAI's frontier o3 model. We tested the model (o3-2025-04-03) with a combination of human users and investigator agents that learned from the human users, to scale the number and types of possible interactions. We also used Docent to analyze the resulting transcripts for surprising behaviors. We surfaced a variety of behaviors that indicate truthfulness issues in o3. Most frequently, o3 hallucinates usage of a code tool and associa
Investigating truthfulness in a pre-release o3 model Neil Chowdhury* , Daniel Johnson , Vincent Huang , Jacob Steinhardt , Sarah Schwettmann* * Correspondence to: neil@transluce.org, sarah@transluce.org Transluce | Published: April 16, 2025 During pre-release testing of OpenAI's o3 model, we found that o3 frequently fabricates actions it took to fulfill user requests, and elaborately justifies the fabrications when confronted by the user. These behaviors generalize to other reasoning models that are widely deployed, such as o3-mini and o1. To dig deeper, we automatically generated hundreds of
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- confessions_paper.pdfcdn.openai.com
- o1 and Reasoning | AndoLogsblog.ando.ai
- Learning to reason with LLMs | OpenAIopenai.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- o1: A Technical Primer — LessWronglesswrong.com
- The o1 System Card Is Not About o1 — LessWronglesswrong.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- o3 — LessWronglesswrong.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com