Digital news platforms use news recommenders as the main instrument to cater to the individual information needs of readers. Despite an increasingly language-diverse online community, in which many Internet users consume news in multiple languages, the majority of news recommendation focuses on major, resource-rich languages, and English in particular. Moreover, nearly all news recommendation efforts assume monolingual news consumption, whereas more and more users tend to consume information in at least two languages. Accordingly, the existing body of work on news recommendation suffers from a lack of publicly available multilingual benchmarks that would catalyze development of news recommenders effective in multilingual settings and for low-resource languages. Aiming to fill this gap, we introduce xMIND, an open, multilingual news recommendation dataset derived from the English MIND dataset using machine translation, covering a set of 14 linguistically and geographically diverse languages, with digital footprints of varying sizes. Using xMIND, we systematically benchmark several state-of-the-art content-based neural news recommenders (NNRs) in both zero-shot (ZS-XLT) and few-shot (FS-XLT) cross-lingual transfer scenarios, considering both monolingual and bilingual news consumption patterns. Our findings reveal that (i) current NNRs, even when based on a multilingual language model, suffer from substantial performance losses under ZS-XLT and that (ii) inclusion of target-language data in FS-XLT training has limited benefits, particularly when combined with a bilingual news consumption. Our findings thus warrant a broader research effort in multilingual and cross-lingual news recommendation. The xMIND dataset is available at https://github.com/andreeaiana/xMIND.
翻译:数字新闻平台通过新闻推荐系统满足读者个性化信息需求。尽管在线社区的语言多样性日益增强,众多互联网用户会使用多种语言消费新闻,但当前绝大多数新闻推荐研究仍聚焦于英语等资源丰富的主流语言。更关键的是,几乎所有新闻推荐工作都假设用户单语种消费新闻,而实际上越来越多用户倾向于使用至少两种语言获取信息。因此,现有新闻推荐研究因缺乏公开的多语言基准数据集,难以推动面向多语言场景及低资源语言的高效新闻推荐系统发展。为填补这一空白,我们基于英语MIND数据集通过机器翻译构建了xMIND——一个开放的多语言新闻推荐数据集,覆盖14种语言及地理分布广泛、数字足迹规模各异的语种。利用xMIND,我们在零样本跨语言迁移(ZS-XLT)和少样本跨语言迁移(FS-XLT)两种场景下,系统评估了多个基于内容的顶尖神经新闻推荐器(NNRs),并同时考虑单语和双语新闻消费模式。实验发现:(i)当前NNR即便基于多语言语言模型,在ZS-XLT场景下仍存在显著性能损失;(ii)在FS-XLT训练中引入目标语言数据带来的性能提升有限——尤其是在双语新闻消费场景下。这些发现表明,多语言及跨语言新闻推荐领域亟需更广泛的研究投入。xMIND数据集已开源:https://github.com/andreeaiana/xMIND。