The convergence of materials science and artificial intelligence has unlocked new opportunities for gathering, analyzing, and generating novel materials sourced from extensive scientific literature. Despite the potential benefits, persistent challenges such as manual annotation, precise extraction, and traceability issues remain. Large language models have emerged as promising solutions to address these obstacles. This paper introduces Functional Materials Knowledge Graph (FMKG), a multidisciplinary materials science knowledge graph. Through the utilization of advanced natural language processing techniques, extracting millions of entities to form triples from a corpus comprising all high-quality research papers published in the last decade. It organizes unstructured information into nine distinct labels, covering Name, Formula, Acronym, Structure/Phase, Properties, Descriptor, Synthesis, Characterization Method, Application, and Domain, seamlessly integrating papers' Digital Object Identifiers. As the latest structured database for functional materials, FMKG acts as a powerful catalyst for expediting the development of functional materials and a fundation for building a more comprehensive material knowledge graph using full paper text. Furthermore, our research lays the groundwork for practical text-mining-based knowledge management systems, not only in intricate materials systems but also applicable to other specialized domains.
翻译:材料科学与人工智能的融合为从大量科学文献中收集、分析和生成新型材料开辟了新机遇。尽管有潜在优势,但手动标注、精确提取和可追溯性等持续挑战依然存在。大语言模型已成为应对这些障碍的有前景解决方案。本文介绍了功能材料知识图谱(FMKG),一个多学科材料科学知识图谱。通过利用先进的自然语言处理技术,从包含过去十年发表的所有高质量研究论文的语料库中提取数百万个实体构成三元组。它将非结构化信息组织成九个不同的标签,涵盖名称、化学式、缩写、结构/相、性质、描述符、合成方法、表征方法、应用和领域,并无缝整合论文的数字对象标识符。作为最新的结构化功能材料数据库,FMKG成为加速功能材料开发的强大催化剂,并为利用全论文文本构建更全面的材料知识图谱奠定基础。此外,我们的研究为基于文本挖掘的实用知识管理系统奠定了基础,不仅适用于复杂的材料系统,也可推广至其他专业领域。