Achieving robust language technologies that can perform well across the world's many languages is a central goal of multilingual NLP. In this work, we take stock of and empirically analyse task performance disparities that exist between multilingual task-oriented dialogue (ToD) systems. We first define new quantitative measures of absolute and relative equivalence in system performance, capturing disparities across languages and within individual languages. Through a series of controlled experiments, we demonstrate that performance disparities depend on a number of factors: the nature of the ToD task at hand, the underlying pretrained language model, the target language, and the amount of ToD annotated data. We empirically prove the existence of the adaptation and intrinsic biases in current ToD systems: e.g., ToD systems trained for Arabic or Turkish using annotated ToD data fully parallel to English ToD data still exhibit diminished ToD task performance. Beyond providing a series of insights into the performance disparities of ToD systems in different languages, our analyses offer practical tips on how to approach ToD data collection and system development for new languages.
翻译:实现能够在全球多种语言中表现良好的稳健语言技术是多语言自然语言处理的核心目标。本研究系统梳理并实证分析了多语言任务导向对话(ToD)系统之间存在的任务性能差异。我们首先定义了系统性能的绝对等价与相对等价的新型量化指标,以捕捉跨语言及语言内部的性能差异。通过一系列受控实验,我们证明性能差异取决于多个因素:ToD任务的具体性质、底层预训练语言模型、目标语言以及ToD标注数据量。我们通过实验证明了当前ToD系统中存在适应性偏差与内在偏差:例如,使用与英语ToD数据完全平行的阿拉伯语或土耳其语标注数据训练的ToD系统,其ToD任务性能依然显著下降。除了提供关于不同语言中ToD系统性能差异的系列见解外,我们的分析还为如何针对新语言开展ToD数据收集与系统开发提供了实用建议。