Inverse design of heterogeneous catalysts remains challenging because catalyst surfaces exhibit substantial structural complexity with coupled surface-adsorbate interactions across a vast chemical space that is difficult to explore efficiently through conventional screening alone. Although machine learning-based high-throughput screening has accelerated catalyst discovery, its efficiency inevitably declines as the search space grows, motivating the development of generative models that can directly construct catalysts with target properties. Here, we present a conditional catalyst generative model based on the Generative Pretrained Transformer architecture with a numerical embedding layer that enables the generation of catalyst structures conditioned on both categorical and continuous properties within a single autoregressive framework. The model was pretrained on 133 million catalyst structures and subsequently fine-tuned on approximately 460,000 optimized structures with associated categorical properties and binding energies for conditional generation. The resulting model achieved 98% structural validity, 95% optimization validity, and high categorical condition fidelity, with a 93 % joint match rate for adsorbate type and composition. For binding energy conditioning, the match rate of approximately 20% represents a four-fold improvement over the baseline training distribution, and the generated distributions shift systematically toward the target values, enabling a 1.5 to 4-fold improvement in screening efficiency for reaction-targeted catalyst discovery without additional fine-tuning. These results show that large-scale autoregressive pre-training, combined with explicit property conditioning, provides a practical route toward controllable catalyst generation and accelerated catalysts discovery.
翻译:多相催化剂的逆向设计仍面临挑战,因为催化剂表面具有显著的结构复杂性,其表面-吸附物相互作用耦合于广阔的化学空间,仅通过传统筛选方法难以高效探索。尽管基于机器学习的高通量筛选加速了催化剂发现,但其效率随搜索空间扩大而不可避免下降,这促使了可直接构建具有目标性质催化剂的生成模型的发展。本文提出了一种基于生成式预训练Transformer架构的条件催化剂生成模型,该模型包含数值嵌入层,能够在单一自回归框架内同时基于分类属性和连续属性生成催化剂结构。模型在1.33亿个催化剂结构上进行了预训练,随后在约46万个优化结构上进行了微调,这些结构关联了分类属性与结合能,用于条件生成。最终模型实现了98%的结构有效性、95%的优化有效性以及高分类条件保真度,其中吸附物类型与组分的联合匹配率达到93%。在结合能条件化方面,约20%的匹配率相较基线训练分布提升了四倍,且生成分布系统性地向目标值偏移,无需额外微调即可将面向反应目标的催化剂发现筛选效率提升1.5至4倍。这些结果表明,大规模自回归预训练结合显式属性条件化为实现可控催化剂生成与加速催化剂发现提供了实用路径。