[2603.18090] MOSS-TTS Technical Report
Abstract:This technical report presents MOSS-TTS, a speech generation foundation model built on a scalable recipe: discrete audio tokens, autoregressive modeling, and large-scale pretraining. Built on MOSS-Audio-Tokenizer, a causal Transformer tokenizer that compresses 24 kHz audio to 12.5 fps with variable-bitrate RVQ and unified semantic-acoustic representations, we release two complementary generators: MOSS-TTS, which emphasizes structural simplicity, scalability, and long-context/control-oriented deployment, and MOSS-TTS-Local-Transformer, which introduces a frame-local autoregressive module for higher modeling efficiency, stronger speaker preservation, and a shorter time to first audio. Across multilingual and open-domain settings, MOSS-TTS supports zero-shot voice cloning, token-level duration control, phoneme-/pinyin-level pronunciation control, smooth code-switching, and stable long-form generation. This report summarizes the design, training recipe, and empirical characteristics of the released models.
OpenMOSS MOSS-TTS Technical Report SII-OpenMOSS Team* Abstract arXiv:2603.18090v2 [cs.SD] 20 Mar 2026 This technical report presents MOSS-TTS, a speech generation foundation model built on a scal- able recipe: discrete audio…
saved by
related reading
- Large Language Diffusion Modelsarxiv.org
- Crossing the uncanny valley of conversational voice | Sesamesesame.com
- Trending Papers - Hugging Facepaperswithcode.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Woosh: A Sound Effects Foundation Modelarxiv.org
- Tincans - Final Technical Reporttincans.ai
- An Open Base Model for Conversational Speechsi.inc
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Voice AI & Voice Agents | An Illustrated Primervoiceaiandvoiceagents.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Free AI Voice Generator & Voice Agents Platform | ElevenLabselevenlabs.io
- Silent speech with ultrasound — Alephalephneuro.com