Research in scientific disciplines evolves, often rapidly, over time with the emergence of novel methodologies and their associated terminologies. While methodologies themselves being conceptual in nature and rather difficult to automatically extract and characterise, in this paper, we seek to develop supervised models for automatic extraction of the names of the various constituents of a methodology, e.g., `R-CNN', `ELMo' etc. The main research challenge for this task is effectively modeling the contexts around these methodology component names in a few-shot or even a zero-shot setting. The main contributions of this paper towards effectively identifying new evolving scientific methodology names are as follows: i) we propose a factored approach to sequence modeling, which leverages a broad-level category information of methodology domains, e.g., `NLP', `RL' etc.; ii) to demonstrate the feasibility of our proposed approach of identifying methodology component names under a practical setting of fast evolving AI literature, we conduct experiments following a simulated chronological setup (newer methodologies not seen during the training process); iii) our experiments demonstrate that the factored approach outperforms state-of-the-art baselines by margins of up to 9.257\% for the methodology extraction task with the few-shot setup.
翻译:科学学科的研究随着新方法论及其相关术语的出现而不断演进,且往往发展迅速。尽管方法论本身具有概念性且难以自动提取与表征,本文旨在开发监督模型以自动提取方法论各组成部分的名称(例如'R-CNN'、'ELMo'等)。本任务的主要研究挑战在于有效建模这些方法论组件名称的上下文环境,尤其是在少样本甚至零样本场景下。本文在有效识别新兴演化学科方法论名称方面的主要贡献如下:i) 我们提出了一种因子化序列建模方法,该方法利用了方法论领域的广义类别信息(例如'NLP'、'RL'等);ii) 为验证所提方法在快速演进AI文献实际场景中识别方法论组件名称的可行性,我们基于模拟时间序列设置(训练过程中未出现的新方法论)开展实验;iii) 实验表明,在少样本设置下的方法论抽取任务中,因子化方法相较最先进基线模型取得了最高达9.257%的性能提升。