Composed image retrieval (CIR) task takes a composed query of image and text, aiming to search relative images for both conditions. Conventional CIR approaches need a training dataset composed of triplets of query image, query text, and target image, which is very expensive to collect. Several recent works have worked on the zero-shot (ZS) CIR paradigm to tackle the issue without using pre-collected triplets. However, the existing ZS-CIR methods show limited backbone scalability and generalizability due to the lack of diversity of the input texts during training. We propose a novel CIR framework, only using language for its training. Our LinCIR (Language-only training for CIR) can be trained only with text datasets by a novel self-supervision named self-masking projection (SMP). We project the text latent embedding to the token embedding space and construct a new text by replacing the keyword tokens of the original text. Then, we let the new and original texts have the same latent embedding vector. With this simple strategy, LinCIR is surprisingly efficient and highly effective; LinCIR with CLIP ViT-G backbone is trained in 48 minutes and shows the best ZS-CIR performances on four different CIR benchmarks, CIRCO, GeneCIS, FashionIQ, and CIRR, even outperforming supervised method on FashionIQ. Code is available at https://github.com/navervision/lincir
翻译:组合图像检索(CIR)任务以图像和文本组成的组合查询为输入,旨在搜索同时满足两种条件的目标图像。传统CIR方法需要由查询图像、查询文本和目标图像三元组构成的训练数据集,此类数据收集成本极高。近年来,部分研究工作致力于零样本(ZS)CIR范式,通过避免使用预收集的三元组来解决该问题。然而,现有ZS-CIR方法因训练阶段输入文本多样性不足,导致骨干网络可扩展性和泛化能力有限。我们提出一种新型CIR框架,仅使用语言进行训练。我们的LinCIR(仅语言训练的CIR)可通过名为自掩蔽投影(SMP)的新型自监督方法,仅使用文本数据集完成训练。我们将文本潜在嵌入投影至token嵌入空间,并通过替换原始文本中的关键词token构建新文本,同时使新文本与原始文本具有相同的潜在嵌入向量。凭借这一简洁策略,LinCIR展现出惊人的高效性与有效性:采用CLIP ViT-G骨干网络的LinCIR仅需48分钟训练,即在CIRCO、GeneCIS、FashionIQ和CIRR四个不同CIR基准测试中取得最佳零样本CIR性能,甚至在FashionIQ上超越了有监督方法。代码开源地址:https://github.com/navervision/lincir