Multimodal summarization (MS) aims to generate a summary from multimodal input. Previous works mainly focus on textual semantic coverage metrics such as ROUGE, which considers the visual content as supplemental data. Therefore, the summary is ineffective to cover the semantics of different modalities. This paper proposes a multi-task cross-modality learning framework (CISum) to improve multimodal semantic coverage by learning the cross-modality interaction in the multimodal article. To obtain the visual semantics, we translate images into visual descriptions based on the correlation with text content. Then, the visual description and text content are fused to generate the textual summary to capture the semantics of the multimodal content, and the most relevant image is selected as the visual summary. Furthermore, we design an automatic multimodal semantics coverage metric to evaluate the performance. Experimental results show that CISum outperforms baselines in multimodal semantics coverage metrics while maintaining the excellent performance of ROUGE and BLEU.
翻译:摘要:多模态摘要旨在从多模态输入生成摘要。现有工作主要关注诸如ROUGE等文本语义覆盖指标,将视觉内容视为辅助数据。因此,摘要无法有效覆盖不同模态的语义。本文提出一种多任务跨模态学习框架(CISum),通过学习多模态文章中的跨模态交互来提升多模态语义覆盖。为获取视觉语义,我们基于图像与文本内容的相关性将其转化为视觉描述。随后,融合视觉描述与文本内容生成文本摘要以捕获多模态内容的语义,并选取最相关的图像作为视觉摘要。此外,我们设计了一种自动多模态语义覆盖指标来评估性能。实验结果表明,CISum在维持ROUGE和BLEU优秀性能的同时,在多模态语义覆盖指标上优于基线方法。