Systematic reviews are crucial for evidence-based medicine as they comprehensively analyse published research findings on specific questions. Conducting such reviews is often resource- and time-intensive, especially in the screening phase, where abstracts of publications are assessed for inclusion in a review. This study investigates the effectiveness of using zero-shot large language models~(LLMs) for automatic screening. We evaluate the effectiveness of eight different LLMs and investigate a calibration technique that uses a predefined recall threshold to determine whether a publication should be included in a systematic review. Our comprehensive evaluation using five standard test collections shows that instruction fine-tuning plays an important role in screening, that calibration renders LLMs practical for achieving a targeted recall, and that combining both with an ensemble of zero-shot models saves significant screening time compared to state-of-the-art approaches.
翻译:系统综述是循证医学的核心工具,通过系统化分析特定问题的现有研究成果来实现循证决策。然而开展此类综述常需耗费大量资源与时间,尤其在筛检阶段——需评估文献摘要以决定是否纳入综述。本研究探究了零样本大语言模型(LLMs)在自动化筛检中的有效性。我们评估了八种不同LLMs的性能,并开发了一种基于预设召回率阈值的校准技术,用于判定文献是否应纳入系统综述。基于五个标准测试集的综合评估表明:指令微调在筛检中具有关键作用,校准技术使LLMs能有效实现目标召回率,而将两者与零样本模型集成相结合,相比现有最优方法可显著缩短筛检时间。