Growing materials data and data-driven informatics drastically promote the discovery and design of materials. While there are significant advancements in data-driven models, the quality of data resources is less studied despite its huge impact on model performance. In this work, we focus on data bias arising from uneven coverage of materials families in existing knowledge. Observing different diversities among crystal systems in common materials databases, we propose an information entropy-based metric for measuring this bias. To mitigate the bias, we develop an entropy-targeted active learning (ET-AL) framework, which guides the acquisition of new data to improve the diversity of underrepresented crystal systems. We demonstrate the capability of ET-AL for bias mitigation and the resulting improvement in downstream machine learning models. This approach is broadly applicable to data-driven materials discovery, including autonomous data acquisition and dataset trimming to reduce bias, as well as data-driven informatics in other scientific domains.
翻译:材料数据的增长与数据驱动信息学极大地促进了材料的发现与设计。尽管数据驱动模型取得了显著进展,但数据资源的质量对模型性能影响巨大,然而相关研究尚不充分。本文聚焦于现有知识中材料家族覆盖不均所导致的数据偏差问题。通过观察常见材料数据库中不同晶系之间的多样性差异,我们提出了一种基于信息熵的度量指标来量化该偏差。为缓解偏差,我们开发了熵目标主动学习(ET-AL)框架,该框架指导新数据的获取以提升代表性不足晶系的多样性。我们验证了ET-AL在偏差缓解方面的能力及其对下游机器学习模型性能的提升效果。该方法可广泛应用于数据驱动的材料发现领域,包括自主数据采集、数据集修剪以减少偏差,以及在其他科学领域的数据驱动信息学应用中。