Typologically diverse benchmarks are increasingly created to track the progress achieved in multilingual NLP. Linguistic diversity of these data sets is typically measured as the number of languages or language families included in the sample, but such measures do not consider structural properties of the included languages. In this paper, we propose assessing linguistic diversity of a data set against a reference language sample as a means of maximising linguistic diversity in the long run. We represent languages as sets of features and apply a version of the Jaccard index suitable for comparing sets of measures. In addition to the features extracted from typological data bases, we propose an automatic text-based measure, which can be used as a means of overcoming the well-known problem of data sparsity in manually collected features. Our diversity score is interpretable in terms of linguistic features and can identify the types of languages that are not represented in a data set. Using our method, we analyse a range of popular multilingual data sets (UD, Bible100, mBERT, XTREME, XGLUE, XNLI, XCOPA, TyDiQA, XQuAD). In addition to ranking these data sets, we find, for example, that (poly)synthetic languages are missing in almost all of them.
翻译:类型多样的基准数据集日益增多,用于追踪多语言自然语言处理(NLP)领域的进展。这些数据集的语种多样性通常通过样本中包含的语言数量或语系数量来衡量,但此类度量并未考虑所包含语言的结构特性。本文提出,以参考语言样本为基准评估数据集的语种多样性,作为长期最大化语种多样性的手段。我们将语言表示为特征集,并采用适用于度量集比较的杰卡德指数变体进行对比。除从类型学数据库中提取的特征外,我们还提出一种基于文本的自动度量方法,可有效解决人工收集特征中常见的数据稀疏问题。我们的多样性分数可通过语言特征进行解释,并能够识别数据集中未被覆盖的语言类型。利用该方法,我们分析了一系列主流的多语言数据集(UD、Bible100、mBERT、XTREME、XGLUE、XNLI、XCOPA、TyDiQA、XQuAD)。除对这些数据集进行排序外,我们进一步发现,例如(多)综合型语言几乎在所有数据集中均缺失。