ResearchPod Summary
This study investigates the alignment between LLM-generated peer reviews and human expert judgments in the context of the ICLR 2026 conference. The researchers compared reviews from three frontier models—OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6—against human reviews for 300 topic-matched submissions. The sample was balanced across three decision categories: rejected, poster, and oral papers. To ensure a controlled comparison, the models were provided with identical instructions, and all decision-related metadata was removed from the papers before evaluation.
The researchers found that while LLMs are capable of broad discrimination, they lack the nuance required for high-stakes conference decision-making. All three models successfully assigned higher ratings to accepted papers than to rejected ones. However, none of the models were able to distinguish between poster and oral papers, a distinction that human reviewers made with statistical significance.
Scoring behavior also varied significantly by provider. Gemini consistently assigned higher ratings than its counterparts, while OpenAI and Claude were more critical, particularly when evaluating high-quality (oral) papers. Furthermore, the models exhibited different thematic priorities: LLMs frequently criticized the lack of baseline comparisons, whereas human reviewers were more likely to raise concerns regarding computational efficiency and resource usage.
These results suggest that while LLMs can serve as useful tools for generating preliminary feedback or summarizing papers, their current judgments should not be considered interchangeable with human expertise. The inability of models to replicate the fine-grained distinctions between presentation categories indicates that they may not fully grasp the subtle criteria that differentiate high-impact research from solid, incremental work. Researchers and conference organizers should treat LLM-generated scores as supplementary rather than definitive.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.