Large-scale vision-language pre-training has shown promising advances on various downstream tasks and achieved significant performance in multi-modal understanding and generation tasks. However, existing methods often perform poorly on image-text matching tasks that require a detailed semantics understanding of the text. Although there have been some works on this problem, they do not sufficiently exploit the structural knowledge present in sentences to enhance multi-modal language representations, which leads to poor performance. In this paper, we present an end-to-end framework Structure-CLIP, which integrates latent detailed semantics from the text to enhance fine-grained semantic representations. Specifically, (1) we use scene graphs in order to pay more attention to the detailed semantic learning in the text and fully explore structured knowledge between fine-grained semantics, and (2) we utilize the knowledge-enhanced framework with the help of the scene graph to make full use of representations of structured knowledge. To verify the effectiveness of our proposed method, we pre-trained our models with the aforementioned approach and conduct experiments on different downstream tasks. Numerical results show that Structure-CLIP can often achieve state-of-the-art performance on both VG-Attribution and VG-Relation datasets. Extensive experiments show its components are effective and its predictions are interpretable, which proves that our proposed method can enhance detailed semantic representation well.
翻译:大规模视觉-语言预训练在各类下游任务中展现出令人瞩目的进展,并在多模态理解与生成任务中取得了显著性能。然而,现有方法在处理需要文本细粒度语义理解的图像-文本匹配任务时往往表现不佳。尽管已有相关研究工作针对此问题展开,但它们未能充分挖掘句子中蕴含的结构知识以增强多模态语言表示,从而导致性能欠佳。本文提出端到端框架Structure-CLIP,通过整合文本中的潜在细粒度语义来增强语义表示的精细程度。具体而言:(1) 我们利用场景图以更加关注文本中的细粒度语义学习,并充分探索细粒度语义之间的结构化知识;(2) 借助场景图构建知识增强框架,以充分利用结构化知识表示。为验证所提方法的有效性,我们采用上述方法预训练模型,并在不同下游任务上进行实验。数值结果表明,Structure-CLIP在VG-Attribution和VG-Relation数据集上均能取得最先进性能。大量实验证实其各组件有效且预测结果具有可解释性,证明本文方法能够有效增强细粒度语义表示。