Node representation learning in a network is an important machine learning technique for encoding relational information in a continuous vector space while preserving the inherent properties and structures of the network. Recently, unsupervised node embedding methods such as DeepWalk, LINE, struc2vec, PTE, UserItem2vec, and RWJBG have emerged from the Skip-gram model and perform better performance in several downstream tasks such as node classification and link prediction than the existing relational models. However, providing post-hoc explanations of Skip-gram-based embeddings remains a challenging problem because of the lack of explanation methods and theoretical studies applicable for embeddings. In this paper, we first show that global explanations to the Skip-gram-based embeddings can be found by computing bridgeness under a spectral cluster-aware local perturbation. Moreover, a novel gradient-based explanation method, which we call GRAPH-wGD, is proposed that allows the top-q global explanations about learned graph embedding vectors more efficiently. Experiments show that the ranking of nodes by scores using GRAPH-wGD is highly correlated with true bridgeness scores. We also observe that the top-q node-level explanations selected by GRAPH-wGD have higher importance scores and produce more changes in class label prediction when perturbed, compared with the nodes selected by recent alternatives, using five real-world graphs.
翻译:网络中的节点表示学习是一种重要的机器学习技术,旨在将关系信息编码至连续向量空间,同时保留网络的内在属性与结构。近年来,基于Skip-gram模型的无监督节点嵌入方法(如DeepWalk、LINE、struc2vec、PTE、UserItem2vec和RWJBG)在节点分类与链接预测等下游任务中展现出优于传统关系模型的性能。然而,由于缺乏适用于嵌入的解释方法与理论研究,为基于Skip-gram的嵌入提供后验解释仍是一个极具挑战性的问题。本文首先证明,通过计算谱聚类感知下的局部扰动桥接度,可以找到基于Skip-gram嵌入的全局解释。此外,我们提出了一种新型基于梯度的解释方法——GRath-wGD,该方法能更高效地获取学习到的图嵌入向量的前q个全局解释。实验表明,使用GRAP-wGD计算的节点分数排序与真实桥接度分数高度相关。我们还观察到,在五个真实世界图上的实验中,与近期替代方法选择的节点相比,由GROUP-wGD选出的前q个节点级解释具有更高的重要性分数,且在受到扰动时对类别标签预测产生更显著的变化。