Recent years have witnessed the rapid development of concept map generation techniques due to their advantages in providing well-structured summarization of knowledge from free texts. Traditional unsupervised methods do not generate task-oriented concept maps, whereas deep generative models require large amounts of training data. In this work, we present GT-D2G (Graph Translation-based Document To Graph), an automatic concept map generation framework that leverages generalized NLP pipelines to derive semantic-rich initial graphs, and translates them into more concise structures under the weak supervision of downstream task labels. The concept maps generated by GT-D2G can provide interpretable summarization of structured knowledge for the input texts, which are demonstrated through human evaluation and case studies on three real-world corpora. Further experiments on the downstream task of document classification show that GT-D2G beats other concept map generation methods. Moreover, we specifically validate the labeling efficiency of GT-D2G in the label-efficient learning setting and the flexibility of generated graph sizes in controlled hyper-parameter studies.
翻译:近年来,概念图生成技术因能够从自由文本中提供结构化的知识摘要而迅速发展。传统的无监督方法无法生成面向任务的概念图,而深度生成模型则需要大量训练数据。本文提出GT-D2G(基于图转换的文档到图)这一自动化概念图生成框架,该框架利用通用NLP流水线获取语义丰富的初始图,并在下游任务标签的弱监督下将其转化为更简洁的结构。通过人工评估及三个真实语料库的案例研究,GT-D2G生成的概念图能够为输入文本提供可解释的结构化知识摘要。在文档分类这一下游任务上的进一步实验表明,GT-D2G优于其他概念图生成方法。此外,我们专门验证了GT-D2G在标签高效学习场景中的标注效率,并通过超参数控制实验验证了生成图大小的灵活性。