Tincans - Final Technical Report
This post describes Tincans' research in pursuit of a real-time AI voice system, largely performed from January to April 2024. We detail four main contributions and release code for each component. We demonstrated the first two components in April 2024, with a demo of sub-500ms AI voice chat on consumer hardware. At the time, it was the fastest public system. The previously undisclosed components are parts of an overall voice chat system capable of real-time duplex chat. The components can be tuned and created cheaply, leveraging existing pretrained models. Constraints breed creativity; we hope that these techniques and code offer inspiration to others. Looking for software engineering roles? Practice interviewing and get a better job with Parahack, the world's best AI tutor! Gazelle is a joint speech-language model which accepts audio natively. We have previously documented it, see our architecture announcement and model release. We also previously opensourced checkpoints and code. He
Tincans - Final Technical Report Tincans Final Technical Report Chris Hua | 2024-09-16 This post describes Tincans' research in pursuit of a real-time AI voice system, largely performed from January to April 2024. We detail four main contributions and release code for each component. Joint speech-language model (Gazelle) Real-time orchestration framework (Gondola) Integrated turn-taking and backchannel prediction model Amortized prefill We demonstrated the first two components in April 2024, with a demo of sub-500ms AI voice chat on consumer hardware. At the time, it was the fastest public sys
Explore this link on the map →saved by
related reading
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Voice AI & Voice Agents | An Illustrated Primervoiceaiandvoiceagents.com
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Labthinkingmachines.ai
- How I built a sub-500ms latency voice agent from scratch | Nick Tikhonovntik.me
- What I've Learned Building Voice Applicationsdeeplearning.ai
- Introducing talkie: a 13B vintage language model from 1930talkie-lm.com
- Silent speech with ultrasound — Alephalephneuro.com
- Advancing voice intelligence with new models in the API | OpenAIopenai.com
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- How Tolan builds voice-first AI with GPT-5.1 | OpenAIopenai.com
- How we solved latency at Vapi - Vapi AI Blogvapi.ai
- GenAI Handbookgenai-handbook.github.io