The use of generative AI to create text descriptions from graphs has mostly focused on knowledge graphs, which connect concepts using facts. In this work we explore the capability of large pretrained language models to generate text from causal graphs, where salient concepts are represented as nodes and causality is represented via directed, typed edges. The causal reasoning encoded in these graphs can support applications as diverse as healthcare or marketing. Using two publicly available causal graph datasets, we empirically investigate the performance of four GPT-3 models under various settings. Our results indicate that while causal text descriptions improve with training data, compared to fact-based graphs, they are harder to generate under zero-shot settings. Results further suggest that users of generative AI can deploy future applications faster since similar performances are obtained when training a model with only a few examples as compared to fine-tuning via a large curated dataset.
翻译:生成式人工智能利用图结构生成文本描述主要集中于知识图谱,这类图谱通过事实连接概念。本研究探索大型预训练语言模型从因果图生成文本的能力,其中关键概念以节点表示,因果关系通过带类型的定向边呈现。这些因果图编码的因果推理可支持医疗保健或市场营销等多样化应用。基于两个公开因果图数据集,我们在不同设置下对四种GPT-3模型进行实证研究。结果表明,虽然因果文本描述随训练数据增加而改善,但与基于事实的图谱相比,其在零样本设置下更难生成。研究进一步表明,生成式人工智能用户可更快部署未来应用,因为仅用少量示例训练模型即可获得与通过大规模精选数据集微调相当的性能。