Research
Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
arXiv:2607.22553v1 Announce Type: cross Abstract: Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we
arXiv:2607.22553v1 Announce Type: cross Abstract: Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review. Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.
Related
- ReviewGuard: Aligning LLM-Assisted Peer Review with Long-Term Scientific Impact
- Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process
- Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI
Source: arXiv cs.AI | 2026-07-28