LLM4SGG: Large Language Model for Weakly Supervised Scene Graph Generation

Weakly-Supervised Scene Graph Generation (WSSGG) research has recently emerged as an alternative to the fully-supervised approach that heavily relies on costly annotations. In this regard, studies on WSSGG have utilized image captions to obtain unlocalized triplets while primarily focusing on grounding the unlocalized triplets over image regions. However, they have overlooked the two issues involved in the triplet formation process from the captions: 1) Semantic over-simplification issue arises when extracting triplets from captions, where fine-grained predicates in captions are undesirably converted into coarse-grained predicates, resulting in a long-tailed predicate distribution, and 2) Low-density scene graph issue arises when aligning the triplets in the caption with entity/predicate classes of interest, where many triplets are discarded and not used in training, leading to insufficient supervision. To tackle the two issues, we propose a new approach, i.e., Large Language Model for weakly-supervised SGG (LLM4SGG), where we mitigate the two issues by leveraging the LLM's in-depth understanding of language and reasoning ability during the extraction of triplets from captions and alignment of entity/predicate classes with target data. To further engage the LLM in these processes, we adopt the idea of Chain-of-Thought and the in-context few-shot learning strategy. To validate the effectiveness of LLM4SGG, we conduct extensive experiments on Visual Genome and GQA datasets, showing significant improvements in both Recall@K and mean Recall@K compared to the state-of-the-art WSSGG methods. A further appeal is that LLM4SGG is data-efficient, enabling effective model training with a small amount of training images.

翻译：弱监督场景图生成（WSSGG）研究近年来作为严重依赖昂贵标注的全监督方法的替代方案而兴起。在此背景下，现有WSSGG研究利用图像标题获取未定位的三元组，并主要关注于在图像区域中对这些未定位三元组进行定位。然而，这些方法忽视了标题中三元组形成过程中涉及的两个问题：1）从标题中提取三元组时出现的语义过度简化问题——标题中的细粒度谓词被不必要地转换为粗粒度谓词，导致谓词分布呈现长尾特征；2）将标题中的三元组与目标实体/谓词类别对齐时产生的低密度场景图问题——大量三元组被丢弃而未用于训练，导致监督信号不足。为解决这两个问题，我们提出新方法LLM4SGG（面向弱监督SGG的大语言模型），通过利用大语言模型对语言的深度理解与推理能力，在从标题中提取三元组以及将实体/谓词类别与目标数据对齐的过程中缓解上述问题。为进一步激发大语言模型在这些过程中的作用，我们借鉴了思维链思想和上下文小样本学习策略。通过在Visual Genome和GQA数据集上的大量实验验证，LLM4SGG在Recall@K和平均Recall@K指标上均显著优于现有最优WSSGG方法。另一个优势在于该方法具有数据高效性，能够使用少量训练图像实现有效模型训练。

相关内容

大语言模型

关注 66

大语言模型是基于海量文本数据训练的深度学习模型。它不仅能够生成自然语言文本，还能够深入理解文本含义，处理各种自然语言任务，如文本摘要、问答、翻译等。2023年，大语言模型及其在人工智能领域的应用已成为全球科技研究的热点，其在规模上的增长尤为引人注目，参数量已从最初的十几亿跃升到如今的一万亿。参数量的提升使得模型能够更加精细地捕捉人类语言微妙之处，更加深入地理解人类语言的复杂性。在过去的一年里，大语言模型在吸纳新知识、分解复杂任务以及图文对齐等多方面都有显著提升。随着技术的不断成熟，它将不断拓展其应用范围，为人类提供更加智能化和个性化的服务，进一步改善人们的生活和生产方式。

O’Reilly报告：知识图谱崛起——面向现代数据集成和数据结构体系，“The Rise of the Knowledge Graph——Toward Modern Data Integration and the Data Fabric Architecture”

专知会员服务

49+阅读 · 2022年2月18日

【CHI2020-微软】解释可解释性:理解数据科学家使用机器学习的可解释性工具，Interpreting Interpretability: Understanding Data Scientists’Use of Interpretability Tools for Machine Learning

专知会员服务

55+阅读 · 2020年3月8日

FlowQA: Grasping Flow in History for Conversational Machine Comprehension

专知会员服务

34+阅读 · 2019年10月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

50+阅读 · 2019年10月17日