Advancing voice intelligence with new models in the API | OpenAI
A new generation of realtime voice models that can reason, translate, and transcribe as people speak. 00:0004:03 We’re introducing three audio models in the API that unlock a new class of voice apps for developers. With these models, developers can build voice experiences that feel more natural, respond more intelligently, and take action in real time: This demo is time-limited. By using it, you agree to OpenAI's Terms and acknowledge our Privacy Policy. Voice is becoming one of the most natural ways for people to use software. It lets someone ask for help while driving, change a travel plan while walking through an airport, get support in their preferred language, or move through a task without stopping to type. But building useful voice products takes more than fast turn-taking or a natural-sounding voice. A voice agent needs to understand what someone means, keep track of context, recover when a request changes, use tools while the conversation continues, and respond in a way that f
May 7, 2026 Product Release Advancing voice intelligence with new models in the API A new generation of realtime voice models that can reason, translate, and transcribe as people speak. Loading… Share We’re introducing three audio models in the API that unlock a new class of voice apps for developers. With these models, developers can build voice experiences that feel more natural, respond more intelligently, and take action in real time: GPT‑Realtime‑2 , our first voice model with GPT‑5‑class reasoning that can handle harder requests and carry the conversation forward naturally. GPT‑Realtime‑
Explore this link on the map →saved by
related reading
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Labthinkingmachines.ai
- Introducing GPT-Live | OpenAIopenai.com
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Voice AI & Voice Agents | An Illustrated Primervoiceaiandvoiceagents.com
- What I've Learned Building Voice Applicationsdeeplearning.ai
- Replicate - Run AI with an APIreplicate.com
- GitHub - openai/whisper: Robust Speech Recognition via Large-Scale Weak Supervision · GitHubgithub.com
- API Overview | OpenAI API Referenceplatform.openai.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Tincans - Final Technical Reporttincans.ai
- How I built a sub-500ms latency voice agent from scratch | Nick Tikhonovntik.me
- Unreal Speech: Cheapest Text-to-Speech APIunrealspeech.com