Graph-structured data, prevalent in domains ranging from social networks to biochemical analysis, serve as the foundation for diverse real-world systems. While graph neural networks demonstrate proficiency in modeling this type of data, their success is often reliant on significant amounts of labeled data, posing a challenge in practical scenarios with limited annotation resources. To tackle this problem, tremendous efforts have been devoted to enhancing graph machine learning performance under low-resource settings by exploring various approaches to minimal supervision. In this paper, we introduce a novel concept of Data-Efficient Graph Learning (DEGL) as a research frontier, and present the first survey that summarizes the current progress of DEGL. We initiate by highlighting the challenges inherent in training models with large labeled data, paving the way for our exploration into DEGL. Next, we systematically review recent advances on this topic from several key aspects, including self-supervised graph learning, semi-supervised graph learning, and few-shot graph learning. Also, we state promising directions for future research, contributing to the evolution of graph machine learning.
翻译:图结构数据广泛存在于从社交网络到生化分析等各个领域,是众多真实世界系统的基础。尽管图神经网络在建模此类数据方面展现了卓越的能力,但其成功往往依赖于大量标注数据,这在实际应用中标注资源有限的场景下构成了挑战。为解决这一问题,研究人员通过探索利用最小监督的各种方法,投入了大量精力来提升低资源环境下的图机器学习性能。本文提出一个新概念——数据高效图学习(DEGL)作为研究前沿,并首次对DEGL的当前进展进行综述。我们首先强调使用大量标注数据训练模型所固有的挑战,为探讨DEGL奠定基础。接着,我们从自监督图学习、半监督图学习和少样图学习等关键方面系统回顾了该主题的最新进展。最后,我们指出了未来有前景的研究方向,以推动图机器学习的发展。