flâneur — a map of the web's best reading

Zero-Shot Text-to-Speech for Vietnamese | OpenReview

openreview.net · 31 words · saved by 1 readers

This paper introduces PhoAudiobook, a newly curated dataset comprising 941 hours of high-quality audio for Vietnamese text-to-speech (TTS). Using PhoAudiobook, we conduct experiments on three leading zero-shot TTS models: VALL-E, VoiceCraft, and XTTS-V2. Our findings demonstrate that PhoAudiobook consistently enhances model performance across various metrics. Moreover, VALL-E and VoiceCraft exhibit superior performance in synthesizing short sentences, highlighting their robustness in handling diverse linguistic contexts. We publicly release PhoAudiobook and the checkpoints of our trained models to facilitate further research and development in Vietnamese TTS. The paper presents an Viet TTS dataset called PhoAudiobook. It clearly stated the why of constructing the dataset and will open to the public later. Experiments are conducted with TTS models to compare the their effectiveness in synthesizing speech. N/A There are no concerns with this submission The paper introduces PhoAudiobook,

Verifying your browser | OpenReview Verifying your browser Complete the check below to continue to OpenReview Please complete the verification above. Have an OpenReview account? Sign in to skip this check.

Explore this link on the map →

related reading