Referring Expression Comprehension (REC) aims to localize the target objects specified by free-form natural language descriptions in images. While state-of-the-art methods achieve impressive performance, they perform a dense perception of images, which incorporates redundant visual regions unrelated to linguistic queries, leading to additional computational overhead. This inspires us to explore a question: can we eliminate linguistic-irrelevant redundant visual regions to improve the efficiency of the model? Existing relevant methods primarily focus on fundamental visual tasks, with limited exploration in vision-language fields. To address this, we propose a coarse-to-fine iterative perception framework, called ScanFormer. It can iteratively exploit the image scale pyramid to extract linguistic-relevant visual patches from top to bottom. In each iteration, irrelevant patches are discarded by our designed informativeness prediction. Furthermore, we propose a patch selection strategy for discarded patches to accelerate inference. Experiments on widely used datasets, namely RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame, verify the effectiveness of our method, which can strike a balance between accuracy and efficiency.
翻译:指代表达式理解(REC)旨在根据图像中自由形式的自然语言描述定位目标对象。虽然最先进的方法取得了令人印象深刻的性能,但它们对图像进行密集感知,其中包含了与语言查询无关的冗余视觉区域,导致额外的计算开销。这启发我们探索一个问题:我们能否消除语言无关的冗余视觉区域以提高模型的效率?现有的相关方法主要关注基础视觉任务,在视觉-语言领域的探索有限。为了解决这个问题,我们提出了一种从粗到细的迭代感知框架,称为ScanFormer。它可以迭代利用图像尺度金字塔从上到下提取语言相关的视觉块。在每次迭代中,不相关的块通过我们设计的信息量预测被丢弃。此外,我们为丢弃的块提出了一种块选择策略以加速推理。在广泛使用的数据集(即RefCOCO、RefCOCO+、RefCOCOg和ReferItGame)上的实验验证了我们方法的有效性,该方法能够在准确性和效率之间取得平衡。