Variable scene layouts and coexisting objects across scenes make indoor scene recognition still a challenging task. Leveraging object information within scenes to enhance the distinguishability of feature representations has emerged as a key approach in this domain. Currently, most object-assisted methods use a separate branch to process object information, combining object and scene features heuristically. However, few of them pay attention to interpretably handle the hidden discriminative knowledge within object information. In this paper, we propose to leverage discriminative object knowledge to enhance scene feature representations. Initially, we capture the object-scene discriminative relationships from a probabilistic perspective, which are transformed into an Inter-Object Discriminative Prototype (IODP). Given the abundant prior knowledge from IODP, we subsequently construct a Discriminative Graph Network (DGN), in which pixel-level scene features are defined as nodes and the discriminative relationships between node features are encoded as edges. DGN aims to incorporate inter-object discriminative knowledge into the image representation through graph convolution. With the proposed IODP and DGN, we obtain state-of-the-art results on several widely used scene datasets, demonstrating the effectiveness of the proposed approach.
翻译:场景布局的多样性及场景间共存目标的存在,使得室内场景识别仍是一项具有挑战性的任务。利用场景中的目标信息增强特征表示的可区分性,已成为该领域的关键方法。目前,多数目标辅助方法采用独立分支处理目标信息,并通过启发式方式融合目标与场景特征。然而,鲜有研究关注如何可解释性地挖掘目标信息中隐藏的判别性知识。本文提出利用目标判别性知识增强场景特征表示。首先,我们从概率角度捕获目标-场景间的判别性关联,并将其转化为目标间判别性原型(IODP)。基于IODP丰富的先验知识,我们进一步构建判别性图网络(DGN),其中像素级场景特征被定义为节点,节点特征间的判别性关系被编码为边。DGN旨在通过图卷积将目标间判别性知识融入图像表示。结合所提出的IODP与DGN,我们在多个广泛使用的场景数据集上取得了最优结果,验证了所提方法的有效性。