Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adversarial attacks remain poorly understood. We present a systematic benchmark of LLM-as-a-Reviewer on 898 papers stratified from NeurIPS and ICLR, evaluating 12 LLMs along three axes: rating calibration, divergence from human reviewers, and resistance to prompt injection embedded via an invisible font-mapping attack. We find that LLMs systematically overrate weaker submissions and diverge from humans in topical emphasis, under-flagging Clarity and over-flagging Reproducibility, while producing reviews two to three times longer with lower lexical diversity and a more standardized vocabulary. Prompt injection remains highly effective. Simple hidden instructions can promote low-scoring papers to acceptance-level ratings in a substantial fraction of cases, with effectiveness varying sharply across model families. While LLMs offer utility in structuring evaluations, their integration into peer review requires safeguards against both intrinsic biases and adversarial risks.
翻译:大型语言模型(LLMs)正越来越多地被用于学术同行评审,但其可靠性、与人类判断的一致性以及对对抗性攻击的鲁棒性仍知之甚少。我们提出了一个系统化的LLM-as-a-Reviewer基准测试,该测试基于从NeurIPS和ICLR中分层选取的898篇论文,沿三个维度评估了12种LLM:评分校准、与人类评审员的分歧,以及对通过隐形字体映射攻击嵌入的提示注入的抵抗力。我们发现,LLMs会系统性地高估较弱的投稿,并在主题重点上与人类存在分歧,低估清晰度而高估可重复性,同时生成的评审意见长度是人类的2到3倍,词汇多样性较低且用词更为标准化。提示注入依然高度有效。简单的隐藏指令能够在相当比例的案例中将低分论文提升到可接受的评分水平,其有效性在不同模型家族间差异显著。尽管LLMs在结构化评估方面具有实用价值,但将其整合到同行评审中仍需针对内在偏见和对抗性风险设立防护措施。