Variable scene layouts and coexisting objects across scenes make indoor scene recognition still a challenging task. Leveraging object information within scenes to enhance the distinguishability of feature representations has emerged as a key approach in this domain. Currently, most object-assisted methods use a separate branch to process object information, combining object and scene features heuristically. However, few of them pay attention to interpretably handle the hidden discriminative knowledge within object information. In this paper, we propose to leverage discriminative object knowledge to enhance scene feature representations. Initially, we capture the object-scene discriminative relationships from a probabilistic perspective, which are transformed into an Inter-Object Discriminative Prototype (IODP). Given the abundant prior knowledge from IODP, we subsequently construct a Discriminative Graph Network (DGN), in which pixel-level scene features are defined as nodes and the discriminative relationships between node features are encoded as edges. DGN aims to incorporate inter-object discriminative knowledge into the image representation through graph convolution and mapping operations (GCN). With the proposed IODP and DGN, we obtain state-of-the-art results on several widely used scene datasets, demonstrating the effectiveness of the proposed approach.
翻译:室内场景布局多变且场景间对象共存,使得室内场景识别仍是一项具有挑战性的任务。利用场景内的对象信息来增强特征表示的判别性,已成为该领域的关键方法。目前,多数对象辅助方法采用独立分支处理对象信息,并启发式地融合对象与场景特征。然而,鲜有方法关注如何可解释地挖掘对象信息中隐藏的判别性知识。本文提出利用判别性对象知识增强场景特征表示。首先,从概率视角捕捉对象-场景判别性关系,并将其转化为对象间判别性原型(IODP)。基于IODP提供的丰富先验知识,进一步构建判别性图网络(DGN),其中像素级场景特征被定义为节点,节点特征间的判别性关系被编码为边。DGN旨在通过图卷积与映射操作将对象间判别性知识融入图像表示。借助所提出的IODP与DGN,我们在多个广泛使用的场景数据集上取得了最先进的成果,验证了该方法的有效性。