Publications

You can also find my articles on my Google Scholar profile.

2026 APSIPA ASC 2026 Published

Vi-SparkRL: Enhancing Tonal Accuracy and Naturalness in Vietnamese TTS via Multi-Objective Rewards

Xuan-Binh Dinh-Thi, Tri-Nhan Do, Tu-Anh Nguyen, Van-An Chu, Dang-Khoa Mac

Recent advances in Large Language Models and neural audio codecs have significantly improved zero-shot Text-to-Speech (TTS) systems; however, maintaining ton...

View paper →
2026 APSIPA ASC 2026 Published

ViNonverbal: The First Vietnamese Nonverbal Speech Dataset via Scalable Pseudo-Labeling

Minh-Ngoc Nguyen, Tu-Anh Nguyen, Tri-Nhan Do, Hoang-Yen Nguyen, Dang-Khoa Mac

Nonverbal vocalizations, such as laughter and fillers, are essential for expressive speech recognition and synthesis, yet they are largely neglected in Vietn...

View paper →
2026 APSIPA ASC 2026 Published

ViInstruct: A Large-Scale Vietnamese Dataset for Instruction-Based Text-to-Speech Synthesis

Thang Ly, Tri-Nhan Do, Tu-Anh Nguyen, Dang-Khoa Mac

Instruction-based text-to-speech (TTS) enables natural-language control of speaking style without reference audio, but no comparable resource exists despite ...

View paper →
2025 FDSE 2025 Withdrawn

Language Identification with Additive Attention BiLSTM and LaBSE Embeddings

Tri-Nhan Do

Language Identification (LangID) plays a critical role in multilingual NLP pipelines, particularly for large-scale web crawling, dataset filtering, and pre-p...

View paper →
2025 MAPR 2025 Published

Adapting WavLM for Vietnamese Speaker Diarization in Real-world Conversations

Tuan-Duy Thang, Van-Huy Nguyen, Tri-Nhan Do, Quoc-Khanh Nguyen, Trung-Kien Phan, Dang-Khoa Mac

While end-to-end neural diarization (EEND) models have achieved state-of-the-art performance, their effectiveness in low-resource languages such as Vietnames...

View paper →
2025 MAPR 2025 Published

Analyzing the Correlation and Impact of Speech Evaluation Metrics on Real-World Speaker Verification and Speech Recognition

Tan-Loc Le, Van-Huy Nguyen, Tri-Nhan Do, Trung-Kien Phan, Dang-Khoa Mac

This study investigates the relationship between speech quality assessment metrics and the performance of downstream tasks such as Automatic Speech Recogniti...

View paper →
2025 MAPR 2025 Published

Unified Acoustic Representation Learning for Vietnamese Speech Classification Tasks

Xuan-Truong Ha, Van-Huy Nguyen, Tri-Nhan Do, Trung-Kien Phan, Dang-Khoa Mac

Analyzing multiple attributes of Vietnamese speech concurrently, such as gender, dialect, and emotion, is an important yet challenging task. Traditional meth...

View paper →
2025 Paper Published

Voice Attacker Leveraging Multi-Head Factorized Attentive Reconstructor and Gradient Reversal for Random Prosody Anonymization

Nhan Tri Do

This is the report for Team 04-SpeechWorld in the First VoicePrivacy Attacker Challenge. The attack methods aimed to verify speakers anonymized by two main a...

View paper →
2024 Paper Published

Enhancing Deepfake Detection: A Study Using WavLM and Advanced RawBoost Augmentation Techniques

Nhan Tri Do, Loi Nguyen Hoang, Phuong Ta Viet, Kien Phan Trung

Automatic Speaker Verification (ASV) systems are increas- ingly vulnerable to sophisticated spoofing attacks, particularly those involving deepfake audio. Th...

View paper →
2023 Detection and Classification of Acoustic Scenes and Events 2023 Published

Sound Event Detection with Soft Labels using Self-Attention Mechanisms for Global Scene Feature Extraction

Nhan Tri-Do, Param Biyani, Zhang Yuxuan, Andrew Koh Jin Jie, Chng Eng Siong

This paper presents our approach to Task 4b of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2023 Challenge, which focuses on Sound ...

View paper →
2023 Journal Published

FastSpeechStyle: Fast, Emotion Controllable and High-Quality Speech Synthesis

Van Thinh Nguyen, Tri-Nhan Do, Hung-Cuong Pham, Tuan Vu Ho, Ngoc-Minh-Khanh Nguyen, Dang-Khoa Mac

The Non-autoregressive text to speech models such as Fastspeech2 can fast synthesize the high quality speech. This model also allows explicit control of the ...

View paper →
2022 Intelligent User Interfaces Conference Published

Vietnamese Speech-based Question Answering over Car Manuals

Tin Duy Vo, Manh Tien Luong, Duong Minh Le, Hieu Minh Tran, Nhan Tri Do, Tuan-Duy Hien Nguyen, Hung Hai Bui, Dat Quoc Nguyen, Dinh Quoc Phung

This paper presents a novel Vietnamese speech-based question answering system QA-CarManual that enables users to ask car-manual-related questions (e.g. how t...

View paper →
2021 Thesis Published

DeepSpeechVC: Voice Cloning Framework with Speech Synthesis and Voice Conversion - Experiment with Speech2Speech techniques to make voice conversion from a small sample of the target speakers

Tri-Nhan Do, Minh-Tri Nguyen under the advice of Prof. Vu Hai Quan, Msc. Xuan-Nam Cao

With motivation of reconstructing voices for people who are mute after an accident or a person who has died and there is a small amount of data about their v...

View paper →
2020 7th NAFOSTED Conference on Information and Computer Science (NICS) Published

Vietnamese Speech Synthesis with End-to-End Model and Text Normalization

Do Tri Nhan, Nguyen Minh Tri, Cao Xuan Nam

Speech synthesis systems are now getting smarter and more natural thanks to the power of deep neural networks. However, each language has a different phonolo...

View paper →
2020 VNUHCM-US_CONF_2020 Published

Speedyspeech Model with Normalization for Faster Vietnamese Speech Synthesis

Do Tri Nhan, Nguyen Minh Tri, Cao Xuan Nam

End to end speech synthesis models have shown better results than traditional methods in terms of intelligence and spontaneity in recent years. However, thes...

View paper →
2020 MediaEval’20 Published

HCMUS at MediaEval 2020: Emotion Classification Using Wavenet Features with SpecAugment and EfficientNet

Tri-Nhan Do, Minh-Tri Nguyen, Hai-Dang Nguyen, Minh-Triet Tran, Xuan-Nam Cao

MediaEval 2020 provided a subset of the MTG-Jamendo dataset, aimed to recognize mood and theme in music. Team HCMUS proposes several solutions to build effic...

View paper →