Large language models (LLMs) have shown promise in automating scientific peer review. However, existing approaches often struggle to generate in-depth reviews supported by concrete evidence. We argue that a key limitation is the lack of flexibility to proactively investigate suspicious parts of a paper based on accumulated evidence, as human reviewers do. In this paper, we explore how to enable an LLM-based review agent to perform such proactive investigation. We find that this can be naturally formulated as a Markov Decision Process (MDP), and propose ProReviewer, a scientific peer review agent that proactively reviews a paper guided by a maintained, structured review log. The structured review log serves as a workspace for the agent to track evidence and intermediate findings collected during review. Experiments show that ProReviewer with an 8B backbone, trained by supervised fine-tuning and optimized by reinforcement learning, achieves the highest average score across five quality dimensions, outperforming prompt-based methods with much larger frontier LLMs by up to 39% and the strongest fine-tuned baseline by 16% relatively. It also attains the highest win rates against baselines in human evaluation.
翻译:大语言模型(LLMs)在自动化科学同行评审方面已展现出潜力。然而,现有方法通常难以生成有具体证据支持的深度评审意见。我们认为,其关键限制在于缺乏像人类审稿人那样基于累积证据主动探究论文可疑部分的灵活性。本文探索了如何使基于LLM的评审智能体执行此类前瞻性探究。我们发现这可以自然地形式化为马尔可夫决策过程(MDP),并提出ProReviewer——一种在维护的、结构化的评审日志指导下主动评审论文的科学同行评审智能体。该结构化评审日志作为智能体追踪评审过程中收集的证据和中间发现的工作空间。实验表明,采用8B主干的ProReviewer,经过监督微调训练并通过强化学习优化,在五个质量维度上取得了最高平均分,相对优于基于提示方法且具有更大前沿LLM的基线模型高达39%,并超过最强微调基线模型16%。在人机评估中,它也取得了对基线模型的最高胜率。