Self-supervised Multi-modal Contrastive Learning (SMCL) remarkably advances modern Vision-Language Pre-training (VLP) models by aligning visual and linguistic modalities. Due to noises in web-harvested text-image pairs, however, scaling up training data volume in SMCL presents considerable obstacles in terms of computational cost and data inefficiency. To improve data efficiency in VLP, we propose Text-aware Image Mixing (TiMix), which integrates mix-based data augmentation techniques into SMCL, yielding significant performance improvements without significantly increasing computational overhead. We provide a theoretical analysis of TiMixfrom a mutual information (MI) perspective, showing that mixed data samples for cross-modal contrastive learning implicitly serve as a regularizer for the contrastive loss. The experimental results demonstrate that TiMix exhibits a comparable performance on downstream tasks, even with a reduced amount of training data and shorter training time, when benchmarked against existing methods. This work empirically and theoretically demonstrates the potential of data mixing for data-efficient and computationally viable VLP, benefiting broader VLP model adoption in practical scenarios.
翻译:自监督多模态对比学习(SMCL)通过对齐视觉与语言模态,显著推动了现代视觉-语言预训练(VLP)模型的发展。然而,由于网络采集的文本-图像对中存在噪声,扩大SMCL训练数据规模在计算成本和数据效率方面带来了巨大挑战。为提升VLP的数据效率,我们提出文本感知图像混合方法(TiMix),将基于混合的数据增强技术融入SMCL,在不显著增加计算开销的前提下实现性能显著提升。我们从互信息(MI)角度对TiMix进行理论分析,证明跨模态对比学习中混合数据样本隐式充当了对比损失的正则化项。实验结果表明,与现有方法相比,TiMix即使在使用更少的训练数据和更短的训练时间时,仍能在下游任务中表现出相当的性能。本研究从实证和理论层面证明了数据混合在实现数据高效与计算可行的VLP中的潜力,有助于推动VLP模型在实际场景中的广泛应用。