Text-to-image retrieval plays a crucial role across various applications, including digital libraries, e-commerce platforms, and multimedia databases, by enabling the search for images using text queries. Despite the advancements in Multimodal Large Language Models (MLLMs), which offer leading-edge performance, their applicability in large-scale, varied, and ambiguous retrieval scenarios is constrained by significant computational demands and the generation of injective embeddings. This paper introduces the Text2Pic Swift framework, tailored for efficient and robust retrieval of images corresponding to extensive textual descriptions in sizable datasets. The framework employs a two-tier approach: the initial Entity-based Ranking (ER) stage addresses the ambiguity inherent in lengthy text queries through a multiple-queries-to-multiple-targets strategy, effectively narrowing down potential candidates for subsequent analysis. Following this, the Summary-based Re-ranking (SR) stage further refines these selections based on concise query summaries. Additionally, we present a novel Decoupling-BEiT-3 encoder, specifically designed to tackle the challenges of ambiguous queries and to facilitate both stages of the retrieval process, thereby significantly improving computational efficiency via vector-based similarity assessments. Our evaluation, conducted on the AToMiC dataset, demonstrates that Text2Pic Swift outperforms current MLLMs by achieving up to an 11.06% increase in Recall@1000, alongside reductions in training and retrieval durations by 68.75% and 99.79%, respectively.
翻译:文本到图像检索在数字图书馆、电子商务平台和多媒体数据库等多种应用中发挥着关键作用,它允许用户通过文本查询来搜索图像。尽管多模态大语言模型(MLLMs)取得了进展并展现了前沿性能,但在大规模、多样化且模糊的检索场景中,其应用受限于巨大的计算需求和单射嵌入的生成。本文提出Text2Pic Swift框架,旨在针对大规模数据集中与长文本描述对应的图像实现高效且鲁棒的检索。该框架采用两层策略:初始的基于实体的排序(ER)阶段通过多查询到多目标策略处理长文本查询中固有的歧义性,有效缩小后续分析的候选范围。随后,基于摘要的重排序(SR)阶段根据简洁的查询摘要进一步优化这些候选结果。此外,我们提出了一种新颖的Decoupling-BEiT-3编码器,专门设计用于解决歧义性查询的挑战,并支持检索过程的两个阶段,从而通过基于向量的相似性评估显著提升计算效率。在AToMiC数据集上的评估表明,Text2Pic Swift在Recall@1000上比当前MLLMs提升了高达11.06%,同时训练和检索时长分别减少了68.75%和99.79%。