The expressive power of graph neural networks is usually measured by comparing how many pairs of graphs or nodes an architecture can possibly distinguish as non-isomorphic to those distinguishable by the $k$-dimensional Weisfeiler-Leman ($k$-WL) test. In this paper, we uncover misalignments between graph machine learning practitioners' conceptualizations of expressive power and $k$-WL through a systematic analysis of the reliability and validity of $k$-WL. We conduct a survey ($n = 18$) of practitioners to surface their conceptualizations of expressive power and their assumptions about $k$-WL. In contrast to practitioners' beliefs, our analysis (which draws from graph theory and benchmark auditing) reveals that $k$-WL does not guarantee isometry, can be irrelevant to real-world graph tasks, and may not promote generalization or trustworthiness. We argue for extensional definitions and measurement of expressive power based on benchmarks. We further contribute guiding questions for constructing such benchmarks, which is critical for graph machine learning practitioners to develop and transparently communicate our understandings of expressive power.
翻译:图神经网络的表达能力通常通过比较其架构能够区分的非同构图或节点对的数量来衡量,并与 $k$ 维 Weisfeiler-Leman($k$-WL)测试可区分的对象进行对比。本文通过系统分析 $k$-WL 的信度和效度,揭示了图机器学习实践者对表达能力的概念化认知与 $k$-WL 之间的错位。我们开展了一项针对实践者的调查($n=18$),以呈现他们对表达能力的概念化理解及其对 $k$-WL 的假设。与实践者认知相反,我们的分析(基于图论和基准审计)表明:$k$-WL 无法保证等距性,可能与现实图任务无关,且未必能促进泛化能力或可信度。我们主张基于基准测试采用外延式定义来测量表达能力,并进一步构建此类基准测试的指导性问题,这对图机器学习实践者发展并透明地交流我们对表达能力的理解至关重要。