In recent years, hashing methods have been popular in the large-scale media search for low storage and strong representation capabilities. To describe objects with similar overall appearance but subtle differences, more and more studies focus on hashing-based fine-grained image retrieval. Existing hashing networks usually generate both local and global features through attention guidance on the same deep activation tensor, which limits the diversity of feature representations. To handle this limitation, we substitute convolutional descriptors for attention-guided features and propose an Attributes Grouping and Mining Hashing (AGMH), which groups and embeds the category-specific visual attributes in multiple descriptors to generate a comprehensive feature representation for efficient fine-grained image retrieval. Specifically, an Attention Dispersion Loss (ADL) is designed to force the descriptors to attend to various local regions and capture diverse subtle details. Moreover, we propose a Stepwise Interactive External Attention (SIEA) to mine critical attributes in each descriptor and construct correlations between fine-grained attributes and objects. The attention mechanism is dedicated to learning discrete attributes, which will not cost additional computations in hash codes generation. Finally, the compact binary codes are learned by preserving pairwise similarities. Experimental results demonstrate that AGMH consistently yields the best performance against state-of-the-art methods on fine-grained benchmark datasets.
翻译:近年来,哈希方法因低存储需求和强表征能力在大规模媒体检索中广受欢迎。为描述外观整体相似但存在细微差异的物体,越来越多研究聚焦于基于哈希的细粒度图像检索。现有哈希网络通常通过注意力机制在同一深度激活张量上生成局部和全局特征,这限制了特征表示的多样性。为解决这一限制,我们用卷积描述子替代注意力引导特征,并提出一种属性分组与挖掘哈希(AGMH),该方法将类别特定的视觉属性分组并嵌入到多个描述子中,生成综合特征表示以实现高效的细粒度图像检索。具体而言,我们设计了一种注意力分散损失(ADL),迫使描述子关注不同的局部区域并捕获多样的细微细节。此外,我们提出一种逐步交互式外部注意力(SIEA),用于挖掘每个描述子中的关键属性并构建细粒度属性与物体之间的关联。该注意力机制专门用于学习离散属性,不会在哈希码生成中增加额外计算。最终,通过保持成对相似性学习紧凑的二进制码。实验结果表明,在细粒度基准数据集上,AGMH相较于现有最先进方法持续取得最优性能。