An accurate description of information is relevant for a range of problems in atomistic modeling, such as sampling methods, detecting rare events, analyzing datasets, or performing uncertainty quantification (UQ) in machine learning (ML)-driven simulations. Although individual methods have been proposed for each of these tasks, they lack a common theoretical background integrating their solutions. Here, we introduce an information theoretical framework that unifies predictions of phase transformations, kinetic events, dataset optimality, and model-free UQ from atomistic simulations, thus bridging materials modeling, ML, and statistical mechanics. We first demonstrate that, for a proposed representation, the information entropy of a distribution of atom-centered environments is a surrogate value for thermodynamic entropy. Using molecular dynamics (MD) simulations, we show that information entropy differences from trajectories can be used to build phase diagrams, identify rare events, and recover classical theories of nucleation. Building on these results, we use this general concept of entropy to quantify information in datasets for ML interatomic potentials (IPs), informing compression, explaining trends in testing errors, and evaluating the efficiency of active learning strategies. Finally, we propose a model-free UQ method for MLIPs using information entropy, showing it reliably detects extrapolation regimes, scales to millions of atoms, and goes beyond model errors. This method is made available as the package QUESTS: Quick Uncertainty and Entropy via STructural Similarity, providing a new unifying theory for data-driven atomistic modeling and combining efforts in ML, first-principles thermodynamics, and simulations.
翻译:信息量化的准确描述对于原子建模中的一系列问题至关重要,例如采样方法、稀有事件检测、数据集分析,以及机器学习驱动模拟中的不确定性量化。尽管针对上述每个任务已有独立方法提出,但它们缺乏整合其解决方案的共同理论基础。本文提出一个信息论框架,能够统一预测原子模拟中的相变、动力学事件、数据集最优性以及模型无关的不确定性量化,从而桥接材料建模、机器学习与统计力学。我们首先证明,在提出的表征下,原子中心环境分布的信息熵是热力学熵的替代值。通过分子动力学模拟,我们展示了轨迹中的信息熵差异可用于构建相图、识别稀有事件,并重现经典成核理论。基于这些结果,我们利用熵的通用概念量化机器学习原子间势数据集中的信息,从而指导数据压缩、解释测试误差趋势,并评估主动学习策略的效率。最后,我们提出一种基于信息熵的机器学习原子间势模型无关不确定性量化方法,证明其能可靠检测外推区域、扩展至百万原子体系,并超越模型误差。该方法已作为QUESTS软件包发布,其全称为“基于结构相似性的快速不确定性与熵计算”,为数据驱动的原子建模提供了新的统一理论,融合了机器学习、第一性原理热力学与模拟方面的研究工作。