We propose a simple strategy for masking image patches during visual-language contrastive learning that improves the quality of the learned representations and the training speed. During each iteration of training, we randomly mask clusters of visually similar image patches, as measured by their raw pixel intensities. This provides an extra learning signal, beyond the contrastive training itself, since it forces a model to predict words for masked visual structures solely from context. It also speeds up training by reducing the amount of data used in each image. We evaluate the effectiveness of our model by pre-training on a number of benchmarks, finding that it outperforms other masking strategies, such as FLIP, on the quality of the learned representation.
翻译:我们提出了一种在视觉-语言对比学习过程中对图像块进行掩码的简单策略,该策略能够提升所学表征的质量及训练速度。在每次训练迭代中,我们根据原始像素强度度量,随机掩码视觉上相似的图像块簇。这种方法在对比训练本身之外提供了额外的学习信号,因为它迫使模型仅根据上下文预测被掩码视觉结构对应的词语。同时,通过减少每张图像中使用的数据量,该方法也加速了训练过程。通过在多个基准数据集上进行预训练评估,我们发现该模型在所学表征质量方面优于其他掩码策略(如FLIP)。