This article presents a new NLP task called structured information inference (SII) to address the complexities of information extraction at the device level in materials science. We accomplished this task by tuning GPT-3 on an existed perovskite solar cell FAIR(Findable, Accessible, Interoperable, Reusable) dataset with 91.8 F1-score and we updated the dataset with all related scientific papers up to now. The produced dataset is formatted and normalized, enabling its direct utilization as input in subsequent data analysis. This feature will enable materials scientists to develop their own models by selecting high-quality review papers within their domain. Furthermore, we designed experiments to predict solar cells' electrical performance and reverse-predict parameters on both material gene and FAIR datesets through LLM. We obtained comparable performance with traditional machine learning methods without feature selection, which demonstrates the potential of large language models to judge materials and design new materials like a materials scientist.
翻译:本文提出一项名为结构化信息推理(SII)的新型自然语言处理任务,旨在解决材料科学中器件层面信息提取的复杂性。我们通过在现有钙钛矿太阳能电池FAIR(可查找、可访问、可互操作、可重用)数据集上微调GPT-3,以91.8%的F1分数完成该任务,并利用截至当前所有相关科学论文更新了该数据集。生成的数据集经格式化与规范化处理,可直接作为后续数据分析的输入。这一特性将使材料科学家能通过筛选各自领域的高质量综述论文来开发专属模型。此外,我们设计实验,借助大型语言模型在材料基因与FAIR数据集上预测太阳能电池的电性能并进行参数反推。在不进行特征选择的情况下,我们获得了与传统机器学习方法相当的预测性能,这证明了大型语言模型能够像材料科学家一样判断材料性能并设计新型材料的潜力。