This paper introduces PMIndiaSum, a new multilingual and massively parallel headline summarization corpus focused on languages in India. Our corpus covers four language families, 14 languages, and the largest to date, 196 language pairs. It provides a testing ground for all cross-lingual pairs. We detail our workflow to construct the corpus, including data acquisition, processing, and quality assurance. Furthermore, we publish benchmarks for monolingual, cross-lingual, and multilingual summarization by fine-tuning, prompting, as well as translate-and-summarize. Experimental results confirm the crucial role of our data in aiding the summarization of Indian texts. Our dataset is publicly available and can be freely modified and re-distributed.
翻译:本文介绍了PMIndiaSum,这是一个专注于印度语言的全新多语言大规模并行标题摘要语料库。该语料库涵盖4个语系、14种语言,并包含迄今为止最多的196种语言对,为所有跨语言对的测试提供了基准平台。我们详细阐述了语料库构建的工作流程,包括数据采集、处理和质量保障。此外,我们通过微调、提示以及翻译-摘要等方法,发布了单语、跨语言和多语言摘要任务的基准测试结果。实验数据证实了本语料库在辅助印度语文本摘要中的关键作用。本数据集已公开提供,并允许自由修改与再分发。