flâneur — a map of the web's best reading

Advancing voice intelligence with new models in the API | OpenAI

openai.com · 1,002 words · saved by 1 readers

A new generation of realtime voice models that can reason, translate, and transcribe as people speak. 00:0004:03 We’re introducing three audio models in the API that unlock a new class of voice apps for developers. With these models, developers can build voice experiences that feel more natural, respond more intelligently, and take action in real time: This demo is time-limited. By using it, you agree to OpenAI's Terms and acknowledge our Privacy Policy. Voice is becoming one of the most natural ways for people to use software. It lets someone ask for help while driving, change a travel plan while walking through an airport, get support in their preferred language, or move through a task without stopping to type. But building useful voice products takes more than fast turn-taking or a natural-sounding voice. A voice agent needs to understand what someone means, keep track of context, recover when a request changes, use tools while the conversation continues, and respond in a way that f

May 7, 2026 Product Release Advancing voice intelligence with new models in the API A new generation of realtime voice models that can reason, translate, and transcribe as people speak. Loading… Share We’re introducing three audio models in the API that unlock a new class of voice apps for developers. With these models, developers can build voice experiences that feel more natural, respond more intelligently, and take action in real time: GPT‑Realtime‑2 , our first voice model with GPT‑5‑class reasoning that can handle harder requests and carry the conversation forward naturally. GPT‑Realtime‑

Explore this link on the map →

saved by

related reading