Although there has been significant interest in applying machine learning techniques to structured data, the expressivity (i.e., a description of what can be learned) of such techniques is still poorly understood. In this paper, we study data transformations based on graph neural networks (GNNs). First, we note that the choice of how a dataset is encoded into a numeric form processable by a GNN can obscure the characterisation of a model's expressivity, and we argue that a canonical encoding provides an appropriate basis. Second, we study the expressivity of monotonic max-sum GNNs, which cover a subclass of GNNs with max and sum aggregation functions. We show that, for each such GNN, one can compute a Datalog program such that applying the GNN to any dataset produces the same facts as a single round of application of the program's rules to the dataset. Monotonic max-sum GNNs can sum an unbounded number of feature vectors which can result in arbitrarily large feature values, whereas rule application requires only a bounded number of constants. Hence, our result shows that the unbounded summation of monotonic max-sum GNNs does not increase their expressive power. Third, we sharpen our result to the subclass of monotonic max GNNs, which use only the max aggregation function, and identify a corresponding class of Datalog programs.
翻译:尽管将机器学习技术应用于结构化数据引起了广泛关注,但这些技术的表达能力(即对可学习内容的描述)仍未被充分理解。本文研究了基于图神经网络(GNN)的数据变换。首先,我们指出,数据集编码为GNN可处理的数值形式的方式可能会模糊模型表达能力的刻画,并论证了规范编码提供了合适的基础。其次,我们研究了单调最大和GNN的表达能力,这类网络覆盖了使用最大和求和聚合函数的GNN子类。我们证明,对每个此类GNN,可以计算出一个Datalog程序,使得将该GNN应用于任何数据集所产生的推理结果,等同于将程序规则对数据集应用一轮所得到的结果。单调最大和GNN可以对无界数量的特征向量求和,导致特征值可能任意大,而规则应用仅需有界数量的常量。因此,我们的结果表明,单调最大和GNN的无界求和并未提升其表达能力。第三,我们将结果进一步精确到仅使用最大聚合函数的单调最大GNN子类,并识别出相应的Datalog程序类别。