This paper studies Large Language Models (LLMs) for structured data--particularly graphs--a crucial data modality that remains underexplored in the LLM literature. We aim to understand when and why the incorporation of structural information inherent in graph data can improve the prediction performance of LLMs on node classification tasks. To address the ``when'' question, we examine a variety of prompting methods for encoding structural information, in settings where textual node features are either rich or scarce. For the ``why'' questions, we probe into two potential contributing factors to the LLM performance: data leakage and homophily. Our exploration of these questions reveals that (i) LLMs can benefit from structural information, especially when textual node features are scarce; (ii) there is no substantial evidence indicating that the performance of LLMs is significantly attributed to data leakage; and (iii) the performance of LLMs on a target node is strongly positively related to the local homophily ratio of the node.
翻译:本文研究大型语言模型(LLMs)在处理结构化数据——特别是图数据——方面的能力,这是LLM文献中尚未充分探索的关键数据模态。我们旨在理解何时以及为何图数据中固有的结构信息能够提升LLM在节点分类任务中的预测性能。针对“何时”问题,我们考察了多种对结构信息进行编码的提示方法,并分别在文本节点特征丰富或稀缺的场景下进行了实验。针对“为何”问题,我们探究了影响LLM性能的两个潜在因素:数据泄露与同质性。通过对这些问题的探索,我们发现:(i)LLM能够从结构信息中获益,尤其是在文本节点特征稀缺的情况下;(ii)没有实质性证据表明LLM的性能显著归因于数据泄露;(iii)LLM在目标节点上的性能与该节点的局部同质性比率呈强正相关。