Recently, commonsense learning has been a hot topic in image-text matching. Although it can describe more graphic correlations, commonsense learning still has some shortcomings: 1) The existing methods are based on triplet semantic similarity measurement loss, which cannot effectively match the intractable negative in image-text sample pairs. 2) The weak generalization ability of the model leads to the poor effect of image and text matching on large-scale datasets. According to these shortcomings. This paper proposes a novel image-text matching model, called Active Mining Sample Pair Semantics image-text matching model (AMSPS). Compared with the single semantic learning mode of the commonsense learning model with triplet loss function, AMSPS is an active learning idea. Firstly, the proposed Adaptive Hierarchical Reinforcement Loss (AHRL) has diversified learning modes. Its active learning mode enables the model to more focus on the intractable negative samples to enhance the discriminating ability. In addition, AMSPS can also adaptively mine more hidden relevant semantic representations from uncommented items, which greatly improves the performance and generalization ability of the model. Experimental results on Flickr30K and MSCOCO universal datasets show that our proposed method is superior to advanced comparison methods.
翻译:最近,常识学习成为图像文本匹配领域的研究热点。然而,虽然常识学习能描述更丰富的图像关联,但现有方法仍存在不足:1)现有方法基于三元组语义相似度度量损失,无法有效处理图像-文本样本对中的难分负样本;2)模型泛化能力弱导致大规模数据集上的图文匹配效果不佳。针对上述问题,本文提出一种新型图像文本匹配模型——主动挖掘样本对语义图像文本匹配模型(AMSPS)。与采用三元组损失函数的常识学习模型的单一语义学习模式不同,AMSPS采用主动学习思想:首先,所提出的自适应分层强化损失(AHRL)具有多样化学习模式,其主动学习模式使模型更关注难分负样本以增强判别能力;此外,AMSPS还能从未标注项中自适应挖掘更多隐式关联语义表征,显著提升模型性能与泛化能力。在Flickr30K和MSCOCO通用数据集上的实验结果表明,本文方法优于现有先进对比方法。