✳flâneur — a map of the web's best reading
Introducing hertz-dev, the first open-source base model for conversational audio generation | blog
si.inc · 1,051 words · saved by 1 readers
We're releasing hertz-dev, an 8.5B parameter autoregressive model for interactive, full-duplex speech generation.
Audio modality is imperative to creating interactive agents that feel natural. Currently the two methods of utilizing audio with generative AI are either diffusion based methods or autoregressive methods . Though diffusion based audio models prove to be good at music generation and small samples, truly interactive audio generation needs to be autoregressive. The largest problems in this field are 1) Getting audio generation that sounds human (ie. non-synthetic as well as handling interruptions well) and 2) Handling realtime generation with two live channels that are both producing information,
Explore this link on the map →related reading
- Redirecting...si.inc
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Labthinkingmachines.ai
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Voice AI & Voice Agents | An Illustrated Primervoiceaiandvoiceagents.com
- Replicate - Run AI with an APIreplicate.com
- Advancing voice intelligence with new models in the API | OpenAIopenai.com
- How I built a sub-500ms latency voice agent from scratch | Nick Tikhonovntik.me
- What I've Learned Building Voice Applicationsdeeplearning.ai
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Unreal Speech: Cheapest Text-to-Speech APIunrealspeech.com
- Woosh: A Sound Effects Foundation Modelarxiv.org
- Introducing GPT-Live | OpenAIopenai.com