Tincans - Final Technical Report
This post describes Tincans' research in pursuit of a real-time AI voice system, largely performed from January to April 2024. We detail four main contributions and release code for each component. We demonstrated the first two components in April 2024, with a demo of sub-500ms AI voice chat on consumer hardware. At the time, it was the fastest public system. The previously undisclosed components are parts of an overall voice chat system capable of real-time duplex chat. The components can be tuned and created cheaply, leveraging existing pretrained models. Constraints breed creativity; we hope that these techniques and code offer inspiration to others. Looking for software engineering roles? Practice interviewing and get a better job with Parahack, the world's best AI tutor! Gazelle is a joint speech-language model which accepts audio natively. We have previously documented it, see our architecture announcement and model release. We also previously opensourced checkpoints and code. He
Tincans - Final Technical Report Tincans Final Technical Report Chris Hua | 2024-09-16 This post describes Tincans' research in pursuit of a real-time AI voice system, largely performed from January to April 2024. We detail four main contributions and release code for each component. Joint speech-language model (Gazelle) Real-time orchestration framework (Gondola) Integrated turn-taking and backchannel prediction model Amortized prefill We demonstrated the first two components in April 2024, with a demo of sub-500ms AI voice chat on consumer hardware. At the time, it was the fastest public sys
saved by
related reading
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Voice AI & Voice Agents | An Illustrated Primervoiceaiandvoiceagents.com
- How I built a sub-500ms latency voice agent from scratch | Nick Tikhonovntik.me
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Labthinkingmachines.ai
- Trending Papers - Hugging Facepaperswithcode.com
- What I've Learned Building Voice Applicationsdeeplearning.ai
- Silent speech with ultrasound — Alephalephneuro.com
- Together AI | The AI Native Cloudtogether.ai
- Free AI Voice Generator & Voice Agents Platform | ElevenLabselevenlabs.io
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- Advancing voice intelligence with new models in the API | OpenAIopenai.com
- Hume AI - The AI toolkit for voice and emotionhume.ai