Compositional generalization is an important ability of language models and has many different manifestations. For data-to-text generation, previous research on this ability is limited to a single manifestation called Systematicity and lacks consideration of large language models (LLMs), which cannot fully cover practical application scenarios. In this work, we propose SPOR, a comprehensive and practical evaluation method for compositional generalization in data-to-text generation. SPOR includes four aspects of manifestations (Systematicity, Productivity, Order invariance, and Rule learnability) and allows high-quality evaluation without additional manual annotations based on existing datasets. We demonstrate SPOR on two different datasets and evaluate some existing language models including LLMs. We find that the models are deficient in various aspects of the evaluation and need further improvement. Our work shows the necessity for comprehensive research on different manifestations of compositional generalization in data-to-text generation and provides a framework for evaluation.
翻译:组合泛化是语言模型的重要能力,具有多种不同表现形式。对于数据到文本生成任务而言,以往关于该能力的研究仅限于单一表现形式(系统性),且缺乏对大型语言模型(LLM)的考量,无法完全覆盖实际应用场景。本文提出SPOR——一种用于数据到文本生成中组合泛化能力的全面实用评估方法。SPOR涵盖四种表现形式(系统性、生产性、顺序不变性和规则可学习性),并能在现有数据集基础上无需额外人工标注即可实现高质量评估。我们在两个不同数据集上展示了SPOR的应用,并对包括LLM在内的现有语言模型进行了评估。结果表明,模型在评估的各个方面均存在不足,亟待进一步改进。本研究揭示了在数据到文本生成中系统研究组合泛化不同表现形式的必要性,并为评估提供了系统性框架。