ResearchPod Summary
Multilingual LLM-based judges are increasingly used to evaluate artificial intelligence models and agents across diverse languages. However, judge behavior is often unstable across prompt languages, leading to systematic score inflation, weak cross-lingual consistency, and concrete backbone-ranking reversals when prompts are localized. This paper treats this instability as a measurement problem. Rather than relying on expensive human annotations to calibrate evaluators, the authors propose a label-free post-hoc calibrator called Consensus-Based Calibration (CBC) that isolates and removes language-backbone interactions from multi-evaluator score matrices.
The multilingual judge score is modeled as a two-way additive layout combining task difficulty, backbone skill, a language-backbone interaction term, a task-language effect, and random noise. When the operational goal is a language-invariant evaluator ranking, the language-backbone interaction term represents the primary source of evaluation bias. The CBC estimator computes this interaction term directly from the data without human labels by applying standard double-centering to the cell-mean score matrix under sum-to-zero constraints. The authors establish an explicit finite-sample concentration bound and prove that the estimator remains unbiased even in the presence of unmodeled task-language interactions, separating its formal properties from classical two-way ANOVA.
Evaluating the method on an expanded eight-language Agent-as-a-Judge benchmark comprising 7,920 judge runs, the authors find that 7 of 15 backbone pairs exhibit statistically significant pairwise rank reversal, and the top-ranked backbone alternates across languages. Applying CBC raises held-out cross-task rank consistency from 0.650 to 0.902. On a separately collected M-RewardBench panel spanning seven languages and 10,500 language-item instances, CBC increases rank consistency and improves panel agreement with public human gold preferences from 68.7% to 76.6%. These results demonstrate that post-hoc double-centering effectively stabilizes multi-evaluator multilingual benchmarks without supervision.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.