We present Dolphin, a novel benchmark that addresses the need for a natural language generation (NLG) evaluation framework dedicated to the wide collection of Arabic languages and varieties. The proposed benchmark encompasses a broad range of 13 different NLG tasks, including dialogue generation, question answering, machine translation, summarization, among others. Dolphin comprises a substantial corpus of 40 diverse and representative public datasets across 50 test splits, carefully curated to reflect real-world scenarios and the linguistic richness of Arabic. It sets a new standard for evaluating the performance and generalization capabilities of Arabic and multilingual models, promising to enable researchers to push the boundaries of current methodologies. We provide an extensive analysis of Dolphin, highlighting its diversity and identifying gaps in current Arabic NLG research. We also offer a public leaderboard that is both interactive and modular and evaluate several models on our benchmark, allowing us to set strong baselines against which researchers can compare.
翻译:我们提出Dolphin,这是一个新颖的基准测试,旨在满足针对广泛阿拉伯语及其变体集合的自然语言生成(NLG)评估框架的需求。该基准涵盖了13项不同的NLG任务,包括对话生成、问答、机器翻译、摘要等。Dolphin包含一个庞大的语料库,由50个测试子集上的40个多样且具有代表性的公开数据集组成,这些数据集经过精心策划,以反映现实场景和阿拉伯语的语言丰富性。它为评估阿拉伯语和多语言模型的性能及泛化能力设立了新标准,有望使研究人员能够突破当前方法的边界。我们对Dolphin进行了广泛分析,突显其多样性并指出了当前阿拉伯语NLG研究中的空白。我们还提供了一个兼具交互性和模块化的公开排行榜,并评估了多个模型在该基准上的表现,从而建立了强基线,研究人员可据此进行比较。