Name disambiguation -- a fundamental problem in online academic systems -- is now facing greater challenges with the increasing growth of research papers. For example, on AMiner, an online academic search platform, about 10% of names own more than 100 authors. Such real-world challenging cases have not been effectively addressed by existing researches due to the small-scale or low-quality datasets that they have used. The development of effective algorithms is further hampered by a variety of tasks and evaluation protocols designed on top of diverse datasets. To this end, we present WhoIsWho owning, a large-scale benchmark with over 1,000,000 papers built using an interactive annotation process, a regular leaderboard with comprehensive tasks, and an easy-to-use toolkit encapsulating the entire pipeline as well as the most powerful features and baseline models for tackling the tasks. Our developed strong baseline has already been deployed online in the AMiner system to enable daily arXiv paper assignments. The public leaderboard is available at http://whoiswho.biendata.xyz/. The toolkit is at https://github.com/THUDM/WhoIsWho. The online demo of daily arXiv paper assignments is at https://na-demo.aminer.cn/arxivpaper.
翻译:姓名消歧——在线学术系统中的基本问题——正面临研究论文持续增长带来的更大挑战。例如,在学术搜索平台AMiner上,约10%的姓名对应超过100位作者。现有研究因采用规模较小或质量较低的数据集,未能有效解决这类现实中的难题。不同数据集上设计的多样化任务与评估协议进一步阻碍了有效算法的发展。为此,我们提出WhoIsWho:一个包含超百万篇论文的大规模基准(通过交互式标注流程构建)、一个附带全面任务的常规排行榜,以及一个易用的工具包(囊括完整流程及最强大的特征与基线模型)。我们开发的强基线已在AMiner系统在线部署,用于每日arXiv论文分配。公开排行榜地址为http://whoiswho.biendata.xyz/,工具包地址为https://github.com/THUDM/WhoIsWho,每日arXiv论文分配的在线演示地址为https://na-demo.aminer.cn/arxivpaper。