ResearchPod Summary
This paper investigates whether LLM-as-a-judge evaluators actually measure the constructs they are intended to assess, rather than merely correlating with surface-level features. The authors formalize construct validity as a two-dimensional profile: invariance (S), the ability to ignore construct-preserving edits, and construct sensitivity (R), the ability to detect construct-changing edits. They test 7 LLM judges across 4 domains using a set of minimal interventions. Crucially, they decompose construct-changing edits into two axes: scope (the breadth of situations a claim covers) and strength (the level of commitment to those situations). Directionality for these edits was established by human annotators, and generation, verification, and judging were performed by disjoint model families to prevent circularity.
This work demonstrates that high agreement scores in LLM evaluation are not evidence of validity. By showing that judges are "blind" to strength-based overreach, the authors provide a mechanism for the paradox where prompting models for accuracy can inadvertently increase overgeneralization. The study provides a framework for researchers to report both invariance and sensitivity, rather than relying on scalar agreement metrics that mask structural failures in evaluation.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.