The last years have witnessed the emergence of a promising self-supervised learning strategy, referred to as masked autoencoding. However, there is a lack of theoretical understanding of how masking matters on graph autoencoders (GAEs). In this work, we present masked graph autoencoder (MaskGAE), a self-supervised learning framework for graph-structured data. Different from standard GAEs, MaskGAE adopts masked graph modeling (MGM) as a principled pretext task - masking a portion of edges and attempting to reconstruct the missing part with partially visible, unmasked graph structure. To understand whether MGM can help GAEs learn better representations, we provide both theoretical and empirical evidence to comprehensively justify the benefits of this pretext task. Theoretically, we establish close connections between GAEs and contrastive learning, showing that MGM significantly improves the self-supervised learning scheme of GAEs. Empirically, we conduct extensive experiments on a variety of graph benchmarks, demonstrating the superiority of MaskGAE over several state-of-the-arts on both link prediction and node classification tasks.
翻译:近年来,一种名为掩码自编码的有前景的自监督学习策略崭露头角。然而,对于掩码机制如何影响图自编码器(GAE)仍缺乏理论层面的理解。本文提出掩码图自编码器(MaskGAE)——一种面向图结构数据的自监督学习框架。与标准GAE不同,MaskGAE采用掩码图建模(MGM)作为基本原则性预训练任务:掩码部分边并尝试利用部分可见的未掩码图结构重建缺失部分。为探究MGM是否有助于GAE学习更优表示,我们从理论与实证层面全面论证了该预训练任务的益处。理论上,我们建立了GAE与对比学习之间的紧密联系,表明MGM能显著提升GAE的自监督学习方案。实证上,我们在多种图基准上开展大量实验,证明了MaskGAE在链接预测和节点分类任务中均优于多种当前最先进方法。