Towards building comprehensive real-world visual perception systems, we propose and study a new problem called panoptic scene graph generation (PVSG). PVSG relates to the existing video scene graph generation (VidSGG) problem, which focuses on temporal interactions between humans and objects grounded with bounding boxes in videos. However, the limitation of bounding boxes in detecting non-rigid objects and backgrounds often causes VidSGG to miss key details crucial for comprehensive video understanding. In contrast, PVSG requires nodes in scene graphs to be grounded by more precise, pixel-level segmentation masks, which facilitate holistic scene understanding. To advance research in this new area, we contribute the PVSG dataset, which consists of 400 videos (289 third-person + 111 egocentric videos) with a total of 150K frames labeled with panoptic segmentation masks as well as fine, temporal scene graphs. We also provide a variety of baseline methods and share useful design practices for future work.
翻译:为构建全面的真实世界视觉感知系统,我们提出并研究了一个名为全景视频场景图生成(PVSG)的新问题。PVSG与现有视频场景图生成(VidSGG)问题相关,后者专注于视频中基于边界框的人与物体之间的时空交互。然而,边界框在检测非刚性物体和背景方面的局限性往往导致VidSGG遗漏关键细节,而这些细节对于全面理解视频至关重要。相比之下,PVSG要求场景图中的节点基于更精确的像素级分割掩码,这有助于实现整体场景理解。为推动这一新领域的研究,我们贡献了PVSG数据集,包含400段视频(289段第三人称视频+111段第一人称视频),总计150K帧,标注了全景分割掩码及精细的时空场景图。我们还提供了多种基线方法,并分享了可供未来工作参考的有用设计实践。