Representation learning is the first step in automating tasks such as research paper recommendation, classification, and retrieval. Due to the accelerating rate of research publication, together with the recognised benefits of interdisciplinary research, systems that facilitate researchers in discovering and understanding relevant works from beyond their immediate school of knowledge are vital. This work explores different methods of research paper representation (or document embedding), to identify those methods that are capable of preserving the interdisciplinary implications of research papers in their embeddings. In addition to evaluating state of the art methods of document embedding in a interdisciplinary citation prediction task, we propose a novel Graph Neural Network architecture designed to preserve the key interdisciplinary implications of research articles in citation network node embeddings. Our proposed method outperforms other GNN-based methods in interdisciplinary citation prediction, without compromising overall citation prediction performance.
翻译:表征学习是自动化研究论文推荐、分类与检索等任务的第一步。鉴于研究论文发表速度日益加快,以及跨学科研究的公认优势,能够帮助研究者发现并理解其直接知识领域之外相关工作的系统至关重要。本研究探讨了不同研究论文表征(或文档嵌入)方法,旨在识别那些能够在嵌入中保留研究论文跨学科含义的方法。除了在跨学科引用预测任务中评估最先进的文档嵌入方法外,我们还提出了一种新颖的图神经网络架构,该架构旨在保留引用网络节点嵌入中研究文章的关键跨学科含义。我们提出的方法在跨学科引用预测方面优于其他基于GNN的方法,且不牺牲整体引用预测性能。