The goal of scene graph generation is to predict a graph from an input image, where nodes correspond to identified and localized objects and edges to their corresponding interaction predicates. Existing methods are trained in a fully supervised manner and focus on message passing mechanisms, loss functions, and/or bias mitigation. In this work we introduce a simple-yet-effective self-supervised relational alignment regularization designed to improve the scene graph generation performance. The proposed alignment is general and can be combined with any existing scene graph generation framework, where it is trained alongside the original model's objective. The alignment is achieved through distillation, where an auxiliary relation prediction branch, that mirrors and shares parameters with the supervised counterpart, is designed. In the auxiliary branch, relational input features are partially masked prior to message passing and predicate prediction. The predictions for masked relations are then aligned with the supervised counterparts after the message passing. We illustrate the effectiveness of this self-supervised relational alignment in conjunction with two scene graph generation architectures, SGTR and Neural Motifs, and show that in both cases we achieve significantly improved performance.
翻译:场景图生成的目标是从输入图像中预测一个图,其中节点对应识别和定位的对象,边对应它们之间的交互谓词。现有方法以全监督方式训练,并专注于消息传递机制、损失函数和/或偏差缓解。本文引入了一种简单而有效的自监督关系对齐正则化方法,旨在提升场景图生成的性能。所提出的对齐方法是通用的,可与任何现有场景图生成框架结合,并随原始模型的训练目标一起训练。该对齐通过蒸馏实现,设计了一个辅助关系预测分支,该分支镜像并共享参数与监督对应分支。在辅助分支中,关系输入特征在消息传递和谓词预测之前被部分掩码。然后,掩码关系的预测在消息传递后与监督对应分支的结果对齐。我们展示了这种自监督关系对齐结合两种场景图生成架构(SGTR 和 Neural Motifs)的有效性,并表明在两种情况下均显著提升了性能。