ViNonverbal: The First Vietnamese Nonverbal Speech Dataset via Scalable Pseudo-Labeling
Published in APSIPA ASC 2026, 2026
Authors: Minh-Ngoc Nguyen, Tu-Anh Nguyen, Tri-Nhan Do, Hoang-Yen Nguyen, Dang-Khoa Mac
Nonverbal vocalizations, such as laughter and fillers, are essential for expressive speech recognition and synthesis, yet they are largely neglected in Vietnamese ASR and TTS systems. We introduce ViNonverbal, a Vietnamese nonverbal speech resource consisting of a manually annotated 500-sample benchmark and an additional 4,693 pseudo-labeled speech segments covering eight nonverbal event types. We fine-tune Qwen3-ASR and Whisper-large-v3 using extended nonverbal tokens to bridge this gap. Experimental results demonstrate that Qwen3-ASR achieves a superior F1-score of 57.21% for nonverbal event recognition, significantly outperforming the Whisper baseline. The model maintains high transcription stability with a WER of 6.16% on VIVOS and 13.61% on the ViNonverbal test set. These findings provide a robust foundation for developing more natural, human-like Vietnamese ASR and TTS systems that can understand and generate expressive nuances.
