Recognizing novel sub-categories with scarce samples is an essential and challenging research topic in computer vision. Existing literature addresses this challenge by employing local-based representation approaches, which may not sufficiently facilitate meaningful object-specific semantic understanding, leading to a reliance on apparent background correlations. Moreover, they primarily rely on high-dimensional local descriptors to construct complex embedding space, potentially limiting the generalization. To address the above challenges, this article proposes a novel model called RSaG for few-shot fine-grained visual recognition. RSaG introduces additional saliency-aware supervision via saliency detection to guide the model toward focusing on the intrinsic discriminative regions. Specifically, RSaG utilizes the saliency detection model to emphasize the critical regions of each sub-category, providing additional object-specific information for fine-grained prediction. RSaG transfers such information with two symmetric branches in a mutual learning paradigm. Furthermore, RSaG exploits inter-regional relationships to enhance the informativeness of the representation and subsequently summarize the highlighted details into contextual embeddings to facilitate the effective transfer, enabling quick generalization to novel sub-categories. The proposed approach is empirically evaluated on three widely used benchmarks, demonstrating its superior performance.
翻译:在计算机视觉中,利用稀缺样本识别新的子类别是一项至关重要且富有挑战性的研究课题。现有文献通过采用基于局部表示的方法来应对这一挑战,但这些方法可能无法充分促进有意义的对象特定语义理解,导致对明显背景相关性的依赖。此外,它们主要依赖高维局部描述子来构建复杂的嵌入空间,可能限制了泛化能力。为解决上述挑战,本文提出了一种名为RSaG的新模型,用于小样本细粒度视觉识别。RSaG通过显著性检测引入额外的显著性感知监督,引导模型关注内在的判别性区域。具体而言,RSaG利用显著性检测模型来强调每个子类别的关键区域,为细粒度预测提供额外的对象特定信息。RSaG通过两个对称分支以互学习范式传递这些信息。此外,RSaG利用区域间关系来增强表示的信息性,随后将突出的细节总结为上下文嵌入,以促进有效传递,从而实现对新的子类别的快速泛化。所提方法在三个广泛使用的基准数据集上进行了实证评估,展示了其优越性能。