We investigate size-induced distribution shifts in graphs and assess their impact on the ability of graph neural networks (GNNs) to generalize to larger graphs relative to the training data. Existing literature presents conflicting conclusions on GNNs' size generalizability, primarily due to disparities in application domains and underlying assumptions concerning size-induced distribution shifts. Motivated by this, we take a data-driven approach: we focus on real biological datasets and seek to characterize the types of size-induced distribution shifts. Diverging from prior approaches, we adopt a spectral perspective and identify that spectrum differences induced by size are related to differences in subgraph patterns (e.g., average cycle lengths). We further find that common GNNs cannot capture these subgraph patterns, resulting in performance decline when testing on larger graphs. Based on these spectral insights, we introduce and compare three model-agnostic strategies aimed at making GNNs aware of important subgraph patterns to enhance their size generalizability: self-supervision, augmentation, and size-insensitive attention. Our empirical results reveal that all strategies enhance GNNs' size generalizability, with simple size-insensitive attention surprisingly emerging as the most effective method. Notably, this strategy substantially enhances graph classification performance on large test graphs, which are 2-10 times larger than the training graphs, resulting in an improvement in F1 scores by up to 8%.
翻译:我们研究了图中由尺寸引起的分布偏移,并评估了其对图神经网络(GNN)泛化到比训练数据更大图的能力的影响。现有文献关于GNN尺寸泛化能力存在相互矛盾的结论,这主要源于应用领域的差异以及关于尺寸引起的分布偏移的不同假设。受此启发,我们采取数据驱动的方法:聚焦于真实生物数据集,旨在刻画由尺寸引起的分布偏移类型。与先前方法不同,我们从频谱视角出发,发现由尺寸引起的频谱差异与子图模式(例如平均环长)的差异密切相关。我们进一步发现,常见的GNN无法捕捉这些子图模式,导致在更大图上测试时性能下降。基于这些频谱见解,我们提出并比较了三种与模型无关的策略,旨在使GNN感知重要的子图模式以增强其尺寸泛化能力:自监督、数据增强和尺寸无关注意力。我们的实证结果表明,所有策略均能提升GNN的尺寸泛化能力,而简单的尺寸无关注意力意外地成为最有效的方法。值得注意的是,该策略显著提升了在大型测试图(其尺寸是训练图的2至10倍)上的图分类性能,F1分数最高提升8%。