Mitigating the climate crisis requires a rapid transition towards lower-carbon energy. Catalyst materials play a crucial role in the electrochemical reactions involved in numerous industrial processes key to this transition, such as renewable energy storage and electrofuel synthesis. To reduce the energy spent on such activities, we must quickly discover more efficient catalysts to drive electrochemical reactions. Machine learning (ML) holds the potential to efficiently model materials properties from large amounts of data, accelerating electrocatalyst design. The Open Catalyst Project OC20 dataset was constructed to that end. However, ML models trained on OC20 are still neither scalable nor accurate enough for practical applications. In this paper, we propose task-specific innovations applicable to most architectures, enhancing both computational efficiency and accuracy. This includes improvements in (1) the graph creation step, (2) atom representations, (3) the energy prediction head, and (4) the force prediction head. We describe these contributions and evaluate them thoroughly on multiple architectures. Overall, our proposed PhAST improvements increase energy MAE by 4 to 42$\%$ while dividing compute time by 3 to 8$\times$ depending on the targeted task/model. PhAST also enables CPU training, leading to 40$\times$ speedups in highly parallelized settings. Python package: \url{https://phast.readthedocs.io}.
翻译:摘要:缓解气候危机需要快速向低碳能源转型。催化剂材料在众多关键工业过程中的电化学反应中扮演着核心角色,这些过程包括可再生能源储存和电燃料合成,对于能源转型至关重要。为了减少此类活动中的能源消耗,我们必须快速发现更高效的催化剂以驱动电化学反应。机器学习(ML)具备从大量数据中高效建模材料特性的潜力,从而加速电催化剂设计。开放催化剂项目OC20数据集正是为此目的而构建。然而,基于OC20训练的ML模型在实际应用中仍缺乏可扩展性和足够的准确性。本文提出了适用于大多数架构的任务特定创新方法,在提升计算效率的同时提高了准确性。这些改进包括:(1)图构建步骤,(2)原子表示方法,(3)能量预测头,以及(4)力预测头。我们详细阐述了这些贡献,并在多种架构上进行了全面评估。总体而言,我们提出的PhAST改进使能量平均绝对误差(MAE)降低了4%至42%,同时根据目标任务/模型的不同,将计算时间缩短至原来的1/3至1/8。此外,PhAST还支持CPU训练,在高并行化场景下可实现40倍加速。Python包:\url{https://phast.readthedocs.io}。