Introducing hertz-dev - Standard Intelligence
For the last few months, the team at Standard Intelligence has been doing research in cross-modality learning. We're excited to announce that we're open-sourcing an early product of this research, an 8.5B, full-duplex, audio-only base model: hertz-dev. Audio modality is imperative to creating interactive agents that feel natural. Currently the two methods of utilizing audio with generative AI are either diffusion based methods or autoregressive methods. Though diffusion based audio models prove to be good at music generation and small samples, truly interactive audio generation needs to be autoregressive. The largest problems in this field are 1) Getting audio generation that sounds human (ie. non-synthetic as well as handling interruptions well) and 2) Handling realtime generation with two live channels that are both producing information, like regular human dialogue. Our model is at the frontier of both of these, natively fitting to the two-speaker format with faster-than-human react
Audio modality is imperative to creating interactive agents that feel natural. Currently the two methods of utilizing audio with generative AI are either diffusion based methods or autoregressive methods. Though diffusion based audio models prove to be good at music generation and small samples, truly interactive audio generation needs to be autoregressive. The largest problems in this field are 1) Getting audio generation that sounds human (ie. non-synthetic as well as handling interruptions well) and 2) Handling realtime generation with two live channels that are both producing…
related reading
- An Open Base Model for Conversational Speech | blogsi.inc
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Labthinkingmachines.ai
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Trending Papers - Hugging Facepaperswithcode.com
- Voice AI & Voice Agents | An Illustrated Primervoiceaiandvoiceagents.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- [2603.18090] MOSS-TTS Technical Reportarxiv.org
- Advancing voice intelligence with new models in the API | OpenAIopenai.com
- Woosh: A Sound Effects Foundation Modelarxiv.org
- Open Yap 1K: 1,000 hours of full-duplex natural conversation, free for commercial usehuggingface.co
- Tiny Interaction Models - Rajan Agarwalrajan.sh
- How I built a sub-500ms latency voice agent from scratch | Nick Tikhonovntik.me