What I've Learned Building Voice Applications
deeplearning.ai · 808 words · saved by 1 readers
The Voice Stack is improving rapidly. Systems that interact with users via speaking and listening will drive many new applications.
Share Loading the Elevenlabs Text to Speech AudioNative Player... Dear friends, The Voice Stack is improving rapidly. Systems that interact with users via speaking and listening will drive many new applications. Over the past year, I’ve been working closely with DeepLearning.AI, AI Fund, and several collaborators on voice-based applications, and I will share best practices I’ve learned in this and future letters. Foundation models that are trained to directly input, and often also directly generate, audio have contributed to this growth, but they are only part of the story. OpenAI’s RealTime A
saved by
related reading
- Voice AI & Voice Agents | An Illustrated Primervoiceaiandvoiceagents.com
- How I built a sub-500ms latency voice agent from scratch | Nick Tikhonovntik.me
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Labthinkingmachines.ai
- How we solved latency at Vapi - Vapi AI Blogvapi.ai
- Advancing voice intelligence with new models in the API | OpenAIopenai.com
- Building Effective AI Agents \ Anthropicanthropic.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Building Effective AI Agents \ Anthropicanthropic.com
- Free AI Voice Generator & Voice Agents Platform | ElevenLabselevenlabs.io
- Tincans - Final Technical Reporttincans.ai
- Hume AI - The AI toolkit for voice and emotionhume.ai