Compositional generalization is an important ability of language models and has many different manifestations. For data-to-text generation, previous research on this ability is limited to a single manifestation called Systematicity and lacks consideration of large language models (LLMs), which cannot fully cover practical application scenarios. In this work, we propose SPOR, a comprehensive and practical evaluation method for compositional generalization in data-to-text generation. SPOR includes four aspects of manifestations (Systematicity, Productivity, Order invariance, and Rule learnability) and allows high-quality evaluation without additional manual annotations based on existing datasets. We demonstrate SPOR on two different datasets and evaluate some existing language models including LLMs. We find that the models are deficient in various aspects of the evaluation and need further improvement. Our work shows the necessity for comprehensive research on different manifestations of compositional generalization in data-to-text generation and provides a framework for evaluation.
翻译:组合泛化是语言模型的重要能力,具有多种不同表现形式。针对数据到文本生成任务,以往关于该能力的研究局限于单一表现形态(系统性,Systematicity),且缺乏对大语言模型(LLMs)的考量,无法全面覆盖实际应用场景。本文提出SPOR——一种面向数据到文本生成中组合泛化能力的综合实用评估方法。该方法涵盖四种表现形态(系统性、生产力、顺序不变性和规则可学习性),并能在现有数据集基础上无需额外人工标注即可实现高质量评估。我们基于两个不同数据集对SPOR进行验证,并评估了包括LLMs在内的若干现有语言模型。研究发现,现有模型在各项评估指标上均存在不足,亟需进一步改进。本工作揭示了在数据到文本生成中对组合泛化能力不同表现形态开展综合研究的必要性,并提供了评估框架。