We introduce HunSum-1: a dataset for Hungarian abstractive summarization, consisting of 1.14M news articles. The dataset is built by collecting, cleaning and deduplicating data from 9 major Hungarian news sites through CommonCrawl. Using this dataset, we build abstractive summarizer models based on huBERT and mT5. We demonstrate the value of the created dataset by performing a quantitative and qualitative analysis on the models' results. The HunSum-1 dataset, all models used in our experiments and our code are available open source.
翻译:我们提出HunSum-1:一个面向匈牙利语抽象式摘要的数据集,包含114万篇新闻文章。该数据集通过CommonCrawl从9个匈牙利主要新闻网站收集、清洗并去重构建而成。基于该数据集,我们构建了基于huBERT和mT5的抽象式摘要模型。通过对模型结果进行定量与定性分析,我们验证了所构建数据集的价值。HunSum-1数据集、实验所用的全部模型及代码均已开源。