Recent developments in LLMs offer new opportunities for assisting authors in improving their work. In this paper, we envision a use case where authors can receive LLM-generated reviews that uncover weak points in the current draft. While initial methods for automated review generation already exist, these methods tend to produce reviews that lack detail, and they do not cover the range of opinions that human reviewers produce. To address this shortcoming, we propose an efficient two-stage review generation framework called Reviewer2. Unlike prior work, this approach explicitly models the distribution of possible aspects that the review may address. We show that this leads to more detailed reviews that better cover the range of aspects that human reviewers identify in the draft. As part of the research, we generate a large-scale review dataset of 27k papers and 99k reviews that we annotate with aspect prompts, which we make available as a resource for future research.
翻译:近期大型语言模型的发展为辅助作者改进其作品提供了新机遇。本文设想了一种应用场景,即作者可获取由大语言模型生成的评论,从而揭示当前草稿中的薄弱环节。尽管已有自动化评论生成方法的初步探索,但现有方法生成的评论缺乏细节,且无法涵盖人类评审员所产生的意见多样性。针对这一不足,我们提出了一种高效的两阶段评论生成框架Reviewer2。与先前工作不同,本方法显式建模了评论可能涉及方面的分布。研究表明,这能生成更详细的评论,更好地覆盖人类评审员在草稿中识别的各类要点。作为研究的一部分,我们构建了包含2.7万篇论文及9.9万条评论的大规模评论数据集,并为其标注了方面提示,将其作为资源供未来研究使用。