Generating natural language text from graph-structured data is essential for conversational information seeking. Semantic triples derived from knowledge graphs can serve as a valuable source for grounding responses from conversational agents by providing a factual basis for the information they communicate. This is especially relevant in the context of large language models, which offer great potential for conversational interaction but are prone to hallucinating, omitting, or producing conflicting information. In this study, we conduct an empirical analysis of conversational large language models in generating natural language text from semantic triples. We compare four large language models of varying sizes with different prompting techniques. Through a series of benchmark experiments on the WebNLG dataset, we analyze the models' performance and identify the most common issues in the generated predictions. Our findings show that the capabilities of large language models in triple verbalization can be significantly improved through few-shot prompting, post-processing, and efficient fine-tuning techniques, particularly for smaller models that exhibit lower zero-shot performance.
翻译:从图结构数据生成自然语言文本对于对话式信息检索至关重要。从知识图谱派生的语义三元组可以成为对话代理响应的有价值基础来源,为其所传达的信息提供事实依据。这一点在大型语言模型(LLMs)的背景下尤为重要,它们虽然在对话交互方面潜力巨大,但容易产生幻觉、遗漏或生成矛盾信息。在本研究中,我们对对话式大型语言模型从语义三元组生成自然语言文本进行了实证分析。我们比较了四种不同规模的大型语言模型,并采用了不同的提示技术。通过在WebNLG数据集上的一系列基准实验,我们分析了这些模型的性能,并识别了生成预测中最常见的问题。我们的研究结果表明,通过少样本提示、后处理和高效微调技术,特别是对于零样本性能较低的较小模型,大型语言模型在三元组语言化方面的能力可以得到显著提升。