Xinzhe Wang, Fei Tao, Jiang Xie, Hong Yu, Ye Wang
4 min
Generating meta-reviews from multiple peer reviews is inherently difficult when reviewers disagree. Traditional approaches often treat all reviewer feedback as equally reliable, which can lead to summaries that either ignore important conflicts or fail to prioritize well-supported arguments. This paper addresses the challenge of synthesizing conflicting evidence by moving beyond uniform aggregation to a reliability-aware approach.
The authors propose a two-level framework that explicitly models the reliability of reviewer evidence. First, the system extracts aspect-level opinions from raw reviews and identifies conflicting sentiments (e.g., positive vs. negative evaluations of novelty). Second, it assigns weights to these opinions using two signals:
By combining these signals, the framework generates a weighted evidence set that allows the meta-review generator to prioritize high-quality, well-supported arguments while still acknowledging diverse perspectives.
Experiments on the ORSum benchmark demonstrate that the proposed method consistently outperforms strong baselines, including those using iterative refinement or sentiment-based consolidation. The framework shows particularly significant performance gains in high-conflict scenarios, where the ability to distinguish between well-justified and vague critiques is most critical. Human evaluations and LLM-as-a-judge assessments confirm that the generated meta-reviews are more coherent, better grounded in the source evidence, and more effective at resolving reviewer disagreements.
This work provides a principled way to handle the subjective and often contradictory nature of academic peer review. By automating the identification and weighting of reliable evidence, the framework helps area chairs and researchers synthesize complex feedback more efficiently, potentially leading to more transparent and evidence-based decision-making in scholarly publishing.
Generating coherent meta-reviews from multiple peer reviews is challenging when reviewer evidence conflicts and varies in reliability. Existing approaches typically formulate meta-review generation as a multi-document summarization task and aggregate reviewer feedback uniformly, making it difficult to determine which opinions should be prioritized under disagreement. In this paper, we study meta-review generation through reliability-aware evidence aggregation. Our framework first extracts aspect-level opinions from peer reviews and identifies conflicting evidence within each aspect. It then estimates opinion-level support and review-level quality to measure evidence reliability. Based on these signals, the framework assigns reliability-aware weights to reviewer feedback, enabling the generator to prioritize better-supported arguments while preserving diverse perspectives. Experiments demonstrate that our method consistently improves meta-review generation over strong baselines on both automatic and human evaluations, with clear gains in conflict recognition and resolution under high-conflict review scenarios. The code and implementation details are publicly available at https://github.com/Wangxz729/reliability-aware-meta-review.
Alex: Then back to the headline. How large is the benefit, and where does it show up?
Sam: On ORSum, reliability-aware weighting beats uniform aggregation on BERTScore and ROUGE-1. The gap widens as reviewer conflict rises, which is what you'd want if the mechanism is filtering noise. But there's a confound. High-conflict cases probably contain richer evidence, so absolute scores there are higher partly for that reason. The relative improvement over baselines is the more informative signal.
Alex: And in low-conflict cases? Could the weighting just get in the way when everyone agrees?
Sam: The paper reports consistent improvements there too. The ablation supports that: when the reliability scores are removed, performance drops. So the load-bearing result is the comparison against uniform aggregation, with the ablation showing the reliability scores are doing the work.
Alex: Which brings us back to the evaluator. If it carries that much weight, we're inheriting whatever biases it has.
Sam: That's the most significant limitation. If the evaluator is biased, those errors propagate into the summary, and the scores are essentially black-box outputs. The authors say so directly. Their proposed direction is human-in-the-loop use, where an area chair could see why a critique was down-weighted, say a tag noting it lacked specificity. That would make the process auditable rather than opaque.
Alex: That seems like the right test of whether this is usable: not only whether the summaries score better, but whether someone can inspect why a point was discounted.
Sam: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Alex: Thanks for listening.