Quantitatively profiling a scholar's scientific impact is important to modern research society. Current practices with bibliometric indicators (e.g., h-index), lists, and networks perform well at scholar ranking, but do not provide structured context for scholar-centric, analytical tasks such as profile reasoning and understanding. This work presents GeneticFlow (GF), a suite of novel graph-based scholar profiles that fulfill three essential requirements: structured-context, scholar-centric, and evolution-rich. We propose a framework to compute GF over large-scale academic data sources with millions of scholars. The framework encompasses a new unsupervised advisor-advisee detection algorithm, a well-engineered citation type classifier using interpretable features, and a fine-tuned graph neural network (GNN) model. Evaluations are conducted on the real-world task of scientific award inference. Experiment outcomes show that the F1 score of best GF profile significantly outperforms alternative methods of impact indicators and bibliometric networks in all the 6 computer science fields considered. Moreover, the core GF profiles, with 63.6%-66.5% nodes and 12.5%-29.9% edges of the full profile, still significantly outrun existing methods in 5 out of 6 fields studied. Visualization of GF profiling result also reveals human explainable patterns for high-impact scholars.
翻译:量化评估学者的科学影响力对现代科研社会至关重要。当前通过文献计量指标(如h指数)、列表和网络的做法虽在学者排序中表现良好,但无法为学者中心的分析任务(如画像推理与理解)提供结构化情境。本研究提出GeneticFlow(GF)系列新型基于图的学者画像,满足三个核心需求:结构化情境、学者中心性及演化丰富性。我们设计了计算框架,可在包含数百万学者的大规模学术数据源上生成GF画像。该框架包含一种新的无监督导师-学生检测算法、基于可解释特征精心设计的引文类型分类器,以及微调的图神经网络(GNN)模型。通过科学奖项推断这一真实任务进行评测,实验结果表明,在全部6个计算机科学领域,最佳GF画像的F1分数显著优于基于影响指标和文献计量网络的替代方法。此外,精简后的核心GF画像(节点保留63.6%-66.5%,边保留12.5%-29.9%)在6个研究领域中的5个仍显著超越现有方法。GF画像结果的可视化还揭示了高影响力学者的可解释模式。