Events serve as fundamental units of occurrence within various contexts. The processing of event semantics in textual information forms the basis of numerous natural language processing (NLP) applications. Recent studies have begun leveraging large language models (LLMs) to address event semantic processing. However, the extent that LLMs can effectively tackle these challenges remains uncertain. Furthermore, the lack of a comprehensive evaluation framework for event semantic processing poses a significant challenge in evaluating these capabilities. In this paper, we propose an overarching framework for event semantic processing, encompassing understanding, reasoning, and prediction, along with their fine-grained aspects. To comprehensively evaluate the event semantic processing abilities of models, we introduce a novel benchmark called EVEVAL. We collect 8 datasets that cover all aspects of event semantic processing. Extensive experiments are conducted on EVEVAL, leading to several noteworthy findings based on the obtained results.
翻译:事件是各类语境中发生的基本单元。文本信息中的事件语义处理构成了众多自然语言处理(NLP)应用的基础。近年来,研究开始利用大语言模型(LLMs)来应对事件语义处理任务。然而,LLMs 在多大程度上能有效解决这些挑战仍不确定。此外,缺乏针对事件语义处理的综合评估框架,给评估这些能力带来了重大挑战。本文提出一个涵盖理解、推理与预测及其细粒度层面的事件语义处理总体框架。为全面评估模型的事件语义处理能力,我们引入名为 EVEVAL 的新基准,并收集覆盖事件语义处理所有维度的 8 个数据集。我们在 EVEVAL 上开展了大量实验,并基于实验结果得出若干值得关注的发现。