The size of the National Aeronautics and Space Administration (NASA) Science Mission Directorate (SMD) is growing exponentially, allowing researchers to make discoveries. However, making discoveries is challenging and time-consuming due to the size of the data catalogs, and as many concepts and data are indirectly connected. This paper proposes a pipeline to generate knowledge graphs (KGs) representing different NASA SMD domains. These KGs can be used as the basis for dataset search engines, saving researchers time and supporting them in finding new connections. We collected textual data and used several modern natural language processing (NLP) methods to create the nodes and the edges of the KGs. We explore the cross-domain connections, discuss our challenges, and provide future directions to inspire researchers working on similar challenges.
翻译:美国国家航空航天局(NASA)科学任务理事会(SMD)的数据规模呈指数级增长,为研究人员提供了新发现的可能性。然而,由于数据目录庞大且许多概念和数据之间存在间接关联,实现这些发现既具挑战性又耗时。本文提出了一种生成知识图谱(KGs)的流水线方法,以表征NASA SMD的不同领域。这些知识图谱可作为数据集搜索引擎的基础,节省研究人员的时间,并支持他们发现新的关联。我们收集了文本数据,并采用多种现代自然语言处理(NLP)方法来构建知识图谱的节点和边。我们探索了跨领域关联,讨论了面临的挑战,并提供了未来方向,以启发面临类似挑战的研究人员。