Text-Based Person Search (TBPS) aims to retrieve pedestrian images using natural language queries. However, existing TBPS models, especially those based on CLIP, struggle with fine-grained understanding due to global representational bias and semantic sparsity inherited from training on short captions. This results in weak fine-grained alignment, exacerbated by the scarcity of region-level annotations. To address this, we propose ROGLE (Robust Global-Local Embedding), a unified framework that overcomes reliance on costly manual annotations through an automated Region-to-Sentence Matching (RSM) strategy. RSM automatically mines pseudo region-sentence pairs for scalable fine-grained supervision. Furthermore, ROGLE employs a multi-granular learning strategy that fuses global contrastive learning with region-level local alignment. We also introduce the P-VLG Benchmark, a large-scale dataset constructed by curating and enriching images from established public benchmarks. It features over 100,000 annotated regions and rich long-form captions, making it the first TBPS benchmark to support both global and local assessment protocols. Extensive experiments show that ROGLE significantly outperforms existing approaches, particularly on challenging long-form queries. Code and the P-VLG benchmark will be made publicly available.
翻译:文本行人搜索(TBPS)旨在利用自然语言查询检索行人图像。然而,现有TBPS模型(尤其是基于CLIP的模型)因继承自短标题训练的全局表示偏置和语义稀疏性,难以实现细粒度理解。这导致弱细粒度对齐问题,且区域级标注的稀缺进一步加剧了该缺陷。为此,我们提出ROGLE(鲁棒全局-局部嵌入)统一框架,通过自动化区域-句子匹配(RSM)策略克服了对昂贵人工标注的依赖。RSM可自动挖掘伪区域-句子对以实现可扩展的细粒度监督。此外,ROGLE采用多粒度学习策略,融合全局对比学习与区域级局部对齐。我们还提出P-VLG基准数据集——通过整理和丰富现有公开基准图像构建的大规模数据集。该数据集包含超过10万个标注区域和丰富长文本描述,成为首个支持全局与局部评估协议的TBPS基准。大量实验表明,ROGLE显著优于现有方法,尤其在具有挑战性的长文本查询场景中表现突出。代码与P-VLG基准将公开发布。