ResearchPod Summary
As Large Audio-Language Models (LALMs) are increasingly used as automated judges for synthetic speech, there is a critical need to understand their reliability beyond simple naturalness scores. This paper introduces ParaPairAudioBench, a diagnostic benchmark consisting of 5,175 audio pairs across five distinct paralinguistic dimensions: Style, Rate, Emphasis, Age, and Gender. By using a pairwise comparison format, the authors evaluate whether these models can accurately distinguish between samples or correctly identify when both samples are equally appropriate (a Tie).
The researchers found that even top-performing LALMs lag behind human judgment by an average of 32 percentage points. A major issue identified is a pervasive calibration failure: models are often unable to abstain from choosing a winner, even when the provided audio samples are indistinguishable or equally poor. While humans frequently select the 'Tie' option in ambiguous scenarios, LALMs almost always force a preference, leading to high error rates in these cases.
Furthermore, the study highlights a significant disparity in how models process different types of information. By comparing 'same-transcript' and 'cross-transcript' conditions, the authors discovered that models rely heavily on lexical (textual) content for style judgments, which causes performance to collapse when the transcripts differ. Conversely, for emphasis, models struggle with localized prosodic cues, often failing to detect word-level stress unless aided by broader sentence-level prosodic context.
This research demonstrates that aggregate naturalness scores are insufficient for evaluating modern speech generation systems. By decomposing evaluation into specific paralinguistic axes, the authors reveal that current LALM judges have systematic blind spots—such as position bias and over-reliance on text—that remain hidden in standard benchmarks. This work provides a necessary framework for developers to diagnose and improve the reliability of LALM-based evaluation pipelines.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.