We introduce ZeroSCROLLS, a zero-shot benchmark for natural language understanding over long texts, which contains only test sets, without training or development data. We adapt six tasks from the SCROLLS benchmark, and add four new datasets, including two novel information fusing tasks, such as aggregating the percentage of positive reviews. Using ZeroSCROLLS, we conduct a comprehensive evaluation of both open-source and closed large language models, finding that Claude outperforms ChatGPT, and that GPT-4 achieves the highest average score. However, there is still room for improvement on multiple open challenges in ZeroSCROLLS, such as aggregation tasks, where models struggle to pass the naive baseline. As the state of the art is a moving target, we invite researchers to evaluate their ideas on the live ZeroSCROLLS leaderboard
翻译:我们提出ZeroSCROLLS,一个针对长文本自然语言理解的零样本基准,仅包含测试集,不提供训练或开发数据。我们基于SCROLLS基准调整了六个任务,并新增了四个数据集,包括两项新颖的信息融合任务(例如聚合正面评论的百分比)。借助ZeroSCROLLS,我们对开源和闭源大型语言模型进行了全面评估,发现Claude的表现优于ChatGPT,而GPT-4取得了最高平均分。然而,在ZeroSCROLLS的多个开放挑战(如聚合任务)中仍有改进空间——当前模型难以超越朴素基线。鉴于领域前沿持续演进,我们邀请研究者通过实时更新的ZeroSCROLLS排行榜验证其创新方法。