Current Scene Graph Generation (SGG) methods explore contextual information to predict relationships among entity pairs. However, due to the diverse visual appearance of numerous possible subject-object combinations, there is a large intra-class variation within each predicate category, e.g., "man-eating-pizza, giraffe-eating-leaf", and the severe inter-class similarity between different classes, e.g., "man-holding-plate, man-eating-pizza", in model's latent space. The above challenges prevent current SGG methods from acquiring robust features for reliable relation prediction. In this paper, we claim that the predicate's category-inherent semantics can serve as class-wise prototypes in the semantic space for relieving the challenges. To the end, we propose the Prototype-based Embedding Network (PE-Net), which models entities/predicates with prototype-aligned compact and distinctive representations and thereby establishes matching between entity pairs and predicates in a common embedding space for relation recognition. Moreover, Prototype-guided Learning (PL) is introduced to help PE-Net efficiently learn such entitypredicate matching, and Prototype Regularization (PR) is devised to relieve the ambiguous entity-predicate matching caused by the predicate's semantic overlap. Extensive experiments demonstrate that our method gains superior relation recognition capability on SGG, achieving new state-of-the-art performances on both Visual Genome and Open Images datasets.
翻译:当前场景图生成方法通过探索上下文信息来预测实体对间的关系。然而,由于大量可能的主语-宾语组合具有多样化的视觉外观,导致模型潜在空间中每个谓词类别内部存在较大差异(例如"人吃披萨"与"长颈鹿吃树叶"),而不同类别之间又存在严重的类间相似性(例如"人拿盘子"与"人吃披萨")。上述挑战阻碍了现有场景图生成方法获取鲁棒特征以实现可靠的关系预测。本文提出,谓词类别固有的语义可作为语义空间中类级原型来缓解这些挑战。为此,我们提出了基于原型的嵌入网络(PE-Net),通过构建与原型对齐的紧凑且具有区分性的实体/谓词表征,在公共嵌入空间中建立实体对与谓词之间的匹配关系用于关系识别。同时引入原型引导学习(PL)帮助PE-Net高效学习这种实体-谓词匹配,并设计了原型正则化(PR)机制缓解因谓词语义重叠导致的模糊实体-谓词匹配。大量实验表明,本方法在场景图生成任务中获得了卓越的关系识别能力,在Visual Genome和Open Images数据集上均取得了新的最先进性能。