Publications
You can also find my articles on my Google Scholar profile.
Vi-SparkRL: Enhancing Tonal Accuracy and Naturalness in Vietnamese TTS via Multi-Objective Rewards
Xuan-Binh Dinh-Thi, Tri-Nhan Do, Tu-Anh Nguyen, Van-An Chu, Dang-Khoa Mac
Recent advances in Large Language Models and neural audio codecs have significantly improved zero-shot Text-to-Speech (TTS) systems; however, maintaining ton...
ViNonverbal: The First Vietnamese Nonverbal Speech Dataset via Scalable Pseudo-Labeling
Minh-Ngoc Nguyen, Tu-Anh Nguyen, Tri-Nhan Do, Hoang-Yen Nguyen, Dang-Khoa Mac
Nonverbal vocalizations, such as laughter and fillers, are essential for expressive speech recognition and synthesis, yet they are largely neglected in Vietn...
ViInstruct: A Large-Scale Vietnamese Dataset for Instruction-Based Text-to-Speech Synthesis
Thang Ly, Tri-Nhan Do, Tu-Anh Nguyen, Dang-Khoa Mac
Instruction-based text-to-speech (TTS) enables natural-language control of speaking style without reference audio, but no comparable resource exists despite ...
Language Identification with Additive Attention BiLSTM and LaBSE Embeddings
Tri-Nhan Do
Language Identification (LangID) plays a critical role in multilingual NLP pipelines, particularly for large-scale web crawling, dataset filtering, and pre-p...
Adapting WavLM for Vietnamese Speaker Diarization in Real-world Conversations
Tuan-Duy Thang, Van-Huy Nguyen, Tri-Nhan Do, Quoc-Khanh Nguyen, Trung-Kien Phan, Dang-Khoa Mac
While end-to-end neural diarization (EEND) models have achieved state-of-the-art performance, their effectiveness in low-resource languages such as Vietnames...
Analyzing the Correlation and Impact of Speech Evaluation Metrics on Real-World Speaker Verification and Speech Recognition
Tan-Loc Le, Van-Huy Nguyen, Tri-Nhan Do, Trung-Kien Phan, Dang-Khoa Mac
This study investigates the relationship between speech quality assessment metrics and the performance of downstream tasks such as Automatic Speech Recogniti...
Unified Acoustic Representation Learning for Vietnamese Speech Classification Tasks
Xuan-Truong Ha, Van-Huy Nguyen, Tri-Nhan Do, Trung-Kien Phan, Dang-Khoa Mac
Analyzing multiple attributes of Vietnamese speech concurrently, such as gender, dialect, and emotion, is an important yet challenging task. Traditional meth...
Voice Attacker Leveraging Multi-Head Factorized Attentive Reconstructor and Gradient Reversal for Random Prosody Anonymization
Nhan Tri Do
This is the report for Team 04-SpeechWorld in the First VoicePrivacy Attacker Challenge. The attack methods aimed to verify speakers anonymized by two main a...
Enhancing Deepfake Detection: A Study Using WavLM and Advanced RawBoost Augmentation Techniques
Nhan Tri Do, Loi Nguyen Hoang, Phuong Ta Viet, Kien Phan Trung
Automatic Speaker Verification (ASV) systems are increas- ingly vulnerable to sophisticated spoofing attacks, particularly those involving deepfake audio. Th...
Sound Event Detection with Soft Labels using Self-Attention Mechanisms for Global Scene Feature Extraction
Nhan Tri-Do, Param Biyani, Zhang Yuxuan, Andrew Koh Jin Jie, Chng Eng Siong
This paper presents our approach to Task 4b of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2023 Challenge, which focuses on Sound ...
FastSpeechStyle: Fast, Emotion Controllable and High-Quality Speech Synthesis
Van Thinh Nguyen, Tri-Nhan Do, Hung-Cuong Pham, Tuan Vu Ho, Ngoc-Minh-Khanh Nguyen, Dang-Khoa Mac
The Non-autoregressive text to speech models such as Fastspeech2 can fast synthesize the high quality speech. This model also allows explicit control of the ...
Vietnamese Speech-based Question Answering over Car Manuals
Tin Duy Vo, Manh Tien Luong, Duong Minh Le, Hieu Minh Tran, Nhan Tri Do, Tuan-Duy Hien Nguyen, Hung Hai Bui, Dat Quoc Nguyen, Dinh Quoc Phung
This paper presents a novel Vietnamese speech-based question answering system QA-CarManual that enables users to ask car-manual-related questions (e.g. how t...
DeepSpeechVC: Voice Cloning Framework with Speech Synthesis and Voice Conversion - Experiment with Speech2Speech techniques to make voice conversion from a small sample of the target speakers
Tri-Nhan Do, Minh-Tri Nguyen under the advice of Prof. Vu Hai Quan, Msc. Xuan-Nam Cao
With motivation of reconstructing voices for people who are mute after an accident or a person who has died and there is a small amount of data about their v...
Vietnamese Speech Synthesis with End-to-End Model and Text Normalization
Do Tri Nhan, Nguyen Minh Tri, Cao Xuan Nam
Speech synthesis systems are now getting smarter and more natural thanks to the power of deep neural networks. However, each language has a different phonolo...
Speedyspeech Model with Normalization for Faster Vietnamese Speech Synthesis
Do Tri Nhan, Nguyen Minh Tri, Cao Xuan Nam
End to end speech synthesis models have shown better results than traditional methods in terms of intelligence and spontaneity in recent years. However, thes...
HCMUS at MediaEval 2020: Emotion Classification Using Wavenet Features with SpecAugment and EfficientNet
Tri-Nhan Do, Minh-Tri Nguyen, Hai-Dang Nguyen, Minh-Triet Tran, Xuan-Nam Cao
MediaEval 2020 provided a subset of the MTG-Jamendo dataset, aimed to recognize mood and theme in music. Team HCMUS proposes several solutions to build effic...