The cross-modal retrieval model leverages the potential of triple loss optimization to learn robust embedding spaces. However, existing methods often train these models in a singular pass, overlooking the distinction between semi-hard and hard triples in the optimization process. The oversight of not distinguishing between semi-hard and hard triples leads to suboptimal model performance. In this paper, we introduce a novel approach rooted in curriculum learning to address this problem. We propose a two-stage training paradigm that guides the model's learning process from semi-hard to hard triplets. In the first stage, the model is trained with a set of semi-hard triplets, starting from a low-loss base. Subsequently, in the second stage, we augment the embeddings using an interpolation technique. This process identifies potential hard negatives, alleviating issues arising from high-loss functions due to a scarcity of hard triples. Our approach then applies hard triplet mining in the augmented embedding space to further optimize the model. Extensive experimental results conducted on two audio-visual datasets show a significant improvement of approximately 9.8% in terms of average Mean Average Precision (MAP) over the current state-of-the-art method, MSNSCA, for the Audio-Visual Cross-Modal Retrieval (AV-CMR) task on the AVE dataset, indicating the effectiveness of our proposed method.
翻译:跨模态检索模型利用三元组损失优化的潜力来学习鲁棒的嵌入空间。然而,现有方法通常以单次训练方式训练这些模型,忽视了优化过程中半困难三元组与困难三元组之间的区分。未区分半困难和困难三元组的疏忽导致模型性能欠优。本文提出一种基于课程学习的新方法来解决这一问题。我们设计了一种两阶段训练范式,引导模型的学习过程从半困难三元组过渡到困难三元组。在第一阶段,模型从低损失基准出发,用一组半困难三元组进行训练。随后在第二阶段,我们采用插值技术增强嵌入,该过程识别潜在的困难负样本,缓解因困难三元组稀缺导致的高损失函数问题。然后,我们的方法在增强后的嵌入空间中应用困难三元组挖掘以进一步优化模型。在两个音视频数据集上进行的大量实验结果表明,在AVE数据集上的音视频跨模态检索任务中,我们的方法相较当前最先进的MSNSCA方法,在平均平均精度均值上实现了约9.8%的显著提升,证明了所提方法的有效性。