Exploring the application of large-scale language models to graph learning is a novel endeavor. However, the vast amount of information inherent in large graphs poses significant challenges to this process. This paper focuses on the link prediction task and introduces LPNL (Link Prediction via Natural Language), a framework based on a large language model designed for scalable link prediction on large-scale heterogeneous graphs.We design novel prompts for link prediction that articulate graph details in natural language. We propose a two-stage sampling pipeline to extract crucial information from large-scale heterogeneous graphs, and a divide-and-conquer strategy to control the input token count within predefined limits, addressing the challenge of overwhelming information. We fine-tune a T5 model based on our self-supervised learning designed for for link prediction. Extensive experiments on a large public heterogeneous graphs demonstrate that LPNL outperforms various advanced baselines, highlighting its remarkable performance in link prediction tasks on large-scale graphs.
翻译:探索将大规模语言模型应用于图学习是一项新颖的研究方向。然而,大型图中包含的海量信息对此过程构成了重大挑战。本文聚焦于链接预测任务,提出LPNL(基于自然语言的链接预测)框架——一种基于大型语言模型的可扩展框架,专为大规模异构图上的链接预测设计。我们设计了新颖的链接预测提示,以自然语言形式表达图的细节信息。通过两阶段采样流程提取大规模异构图中的关键信息,并采用分治策略将输入令牌数量控制在预设范围内,从而应对信息过载问题。基于专为链接预测设计的自监督学习方法,我们对T5模型进行了微调。在大型公开异构图上的广泛实验表明,LPNL在各项先进基线中表现优异,凸显了其在大规模图链接预测任务中的卓越性能。