Panoptic narrative grounding (PNG) aims to segment things and stuff objects in an image described by noun phrases of a narrative caption. As a multimodal task, an essential aspect of PNG is the visual-linguistic interaction between image and caption. The previous two-stage method aggregates visual contexts from offline-generated mask proposals to phrase features, which tend to be noisy and fragmentary. The recent one-stage method aggregates only pixel contexts from image features to phrase features, which may incur semantic misalignment due to lacking object priors. To realize more comprehensive visual-linguistic interaction, we propose to enrich phrases with coupled pixel and object contexts by designing a Phrase-Pixel-Object Transformer Decoder (PPO-TD), where both fine-grained part details and coarse-grained entity clues are aggregated to phrase features. In addition, we also propose a PhraseObject Contrastive Loss (POCL) to pull closer the matched phrase-object pairs and push away unmatched ones for aggregating more precise object contexts from more phrase-relevant object tokens. Extensive experiments on the PNG benchmark show our method achieves new state-of-the-art performance with large margins.
翻译:全景叙事定位(Panoptic Narrative Grounding, PNG)旨在根据叙述性标题中的名词短语,分割图像中的物体与填充物。作为一项多模态任务,PNG的关键在于图像与标题之间的视觉-语言交互。以往的两阶段方法将离线生成的掩码提议中的视觉上下文聚合到短语特征中,但这类特征往往噪声较大且不完整。近期的一阶段方法仅从图像特征中聚合像素上下文至短语特征,由于缺乏对象先验知识,可能导致语义对齐偏差。为实现更全面的视觉-语言交互,我们提出通过设计短语-像素-对象变换器解码器(Phrase-Pixel-Object Transformer Decoder, PPO-TD)来利用耦合的像素与对象上下文增强短语,从而将细粒度部分细节与粗粒度实体线索聚合到短语特征中。此外,我们提出短语-对象对比损失(Phrase-Object Contrastive Loss, POCL),以拉近匹配的短语-对象对并推开不匹配的对,从而从更具短语相关性的对象令牌中聚合更精确的对象上下文。在PNG基准上的大量实验表明,我们的方法以较大优势实现了新的最优性能。