ResearchPod Summary
Speaker recognition performance is heavily dependent on the availability of large-scale, diverse training data. While English and Chinese have benefited from massive corpora like VoxCeleb and CN-Celeb, Vietnamese remains under-resourced. Existing Vietnamese datasets often rely on visual cues (face detection) to link speech to identities, which limits data collection to on-camera recordings and excludes audio-only sources like podcasts or radio. This paper asks whether a face-independent pipeline can effectively construct a large-scale, diverse Vietnamese speaker dataset.
The authors propose a new construction pipeline that replaces visual supervision with textual reasoning. The process involves three main stages:
VieSpeaker provides 902 hours of speech from 4,715 speakers, significantly expanding the scale of available Vietnamese resources. Experimental results demonstrate that models trained on VieSpeaker outperform those trained on previous Vietnamese datasets (Vietnam-Celeb and VoxVietnam) in both standalone training and as a pretraining resource. Specifically, VieSpeaker-pretrained models show improved robustness and generalization, achieving lower Equal Error Rates (EER) on challenging cross-session evaluation protocols.
This work demonstrates that high-quality, large-scale speaker recognition datasets can be built without visual dependency. By tapping into the vast amount of audio-only content available online, this approach provides a scalable path to address data scarcity in under-resourced languages. The release of VieSpeaker offers the research community a more robust benchmark for developing Vietnamese speaker recognition systems.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.