ResearchPod Summary
Conventional Mean Opinion Score (MOS) prediction models are designed to assess global speech naturalness but often fail to detect localized prosodic errors. In Japanese, where pitch-accent placement is critical for lexical meaning, these models are frequently insensitive to subtle accent nucleus errors. The authors investigate whether a specialized model can be trained to explicitly identify and quantify these localized pitch-accent degradations.
To overcome the lack of large-scale datasets with pitch-accent error labels, the authors constructed a controlled dataset using an accent-controllable text-to-speech (TTS) system. They generated speech samples with varying degrees of accent-nucleus errors and assigned them pseudo-quality scores based on the error rate. The proposed model, PASQA, utilizes a self-supervised learning (SSL) backbone augmented with four key strategies:
Experimental results demonstrate that conventional MOS models perform near chance levels when evaluating pitch-accent severity. In contrast, PASQA achieves high ordering accuracy and strong correlation with human listeners on both seen and unseen speakers. Ablation studies confirm that each architectural component—particularly the mora-conditioned fusion and the auxiliary frame-level head—contributes significantly to the model's ability to detect mild errors. Furthermore, PASQA maintains robust performance when evaluated on out-of-domain synthetic speech, proving its utility as a reliable tool for evaluating the prosodic quality of modern TTS systems.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.