ViInstruct: A Large-Scale Vietnamese Dataset for Instruction-Based Text-to-Speech Synthesis

Published in APSIPA ASC 2026, 2026

Authors: Thang Ly, Tri-Nhan Do, Tu-Anh Nguyen, Dang-Khoa Mac

Download Paper

Instruction-based text-to-speech (TTS) enables natural-language control of speaking style without reference audio, but no comparable resource exists despite the language's tonal prosody and substantial regional variation. This paper introduces ViInstruct, the first large-scale Vietnamese dataset for instruction-based TTS, comprising approximately 1,718 hours of speech and 1,224,945 utterance-instruction pairs collected from six corpora spanning general speech, regional accent, and emotion settings. We develop an automated pipeline that extracts acoustic features, preserves or infers speaker attributes, maps them to Vietnamese descriptors, and generates one natural-language instruction for each utterance with a large language model. We also provide a standardized benchmark, comprising two evaluation subsets and two baseline systems, to support reproducible evaluation of controllable Vietnamese TTS in a low-resource tonal language.