Many score-based active learning methods have been successfully applied to graph-structured data, aiming to reduce the number of labels and achieve better performance of graph neural networks based on predefined score functions. However, these algorithms struggle to learn policy distributions that are proportional to rewards and have limited exploration capabilities. In this paper, we innovatively formulate the graph active learning problem as a generative process, named GFlowGNN, which generates various samples through sequential actions with probabilities precisely proportional to a predefined reward function. Furthermore, we propose the concept of flow nodes and flow features to efficiently model graphs as flows based on generative flow networks, where the policy network is trained with specially designed rewards. Extensive experiments on real datasets show that the proposed approach has good exploration capability and transferability, outperforming various state-of-the-art methods.
翻译:许多基于分数的主动学习方法已成功应用于图结构数据,旨在通过预定义评分函数减少标签数量并提升图神经网络性能。然而,这些算法难以学习与奖励成比例的策略分布,且探索能力有限。本文创新性地将图主动学习问题形式化为生成过程(命名为GFlowGNN),通过序列化动作生成多样样本,其概率与预定义奖励函数精确成比例。进一步提出流节点与流特征概念,基于生成流网络将图高效建模为流,并采用专门设计的奖励训练策略网络。在真实数据集上的大量实验表明,该方法具备良好的探索能力与迁移性,性能优于多种现有先进方法。