Pretrained language models (PLMs) have made remarkable progress in table-to-text generation tasks. However, the lack of domain-specific knowledge makes it challenging to bridge the topological gap between tabular data and text, especially in real-world applications with limited resources. To mitigate the limitation of insufficient labeled data, we propose a novel framework: Adapt-Knowledge-to-Generate (AKG). The core insight of AKG is to adapt unlabeled domain-specific knowledge into the model, which brings at least three benefits: (1) it injects representation of normal table-related descriptions to bridge the topological gap between tabular data and texts; (2) it enables us to use large amounts of unlabeled domain-specific knowledge fully, which can alleviate the PLMs' inherent shortcomings of lacking domain knowledge; (3) it allows us to design various tasks to employ the domain-specific knowledge. Extensive experiments and analyses are conducted on three open-domain, few-shot natural language generation (NLG) data sets: Humans, Songs, and Books. Compared to previous state-of-the-art approaches, our model achieves superior performance in terms of both fluency and accuracy as judged by human and automatic evaluations.
翻译:预训练语言模型(PLMs)在表格到文本生成任务中取得了显著进展。然而,由于缺乏领域特定知识,在资源有限的实际应用中,弥补表格数据与文本之间的拓扑差距仍面临挑战。为缓解标注数据不足的局限性,我们提出了一种新框架:Adapt-Knowledge-to-Generate(AKG)。AKG的核心思想是将未标注的领域特定知识自适应地融入模型,这至少带来三个优势:(1)注入与表格相关描述的常规表征,以弥合表格数据与文本之间的拓扑差距;(2)能够充分利用大规模的未标注领域特定知识,从而缓解PLMs固有缺乏领域知识的不足;(3)允许设计多种任务以运用领域特定知识。我们在三个开放域少样本自然语言生成(NLG)数据集(Humans、Songs和Books)上进行了大量实验和分析。与先前最先进的方法相比,我们的模型在人类和自动评估中,在流畅性和准确性方面均取得了更优性能。