Deep generative diffusion models are a promising avenue for de novo 3D molecular design in material science and drug discovery. However, their utility is still constrained by suboptimal performance with large molecular structures and limited training data. Addressing this gap, we explore the design space of E(3) equivariant diffusion models, focusing on previously blank spots. Our extensive comparative analysis evaluates the interplay between continuous and discrete state spaces. Out of this investigation, we introduce the EQGAT-diff model, which consistently surpasses the performance of established models on the QM9 and GEOM-Drugs datasets by a large margin. Distinctively, EQGAT-diff takes continuous atomic positions while chemical elements and bond types are categorical and employ a time-dependent loss weighting that significantly increases training convergence and the quality of generated samples. To further strengthen the applicability of diffusion models to limited training data, we examine the transferability of EQGAT-diff trained on the large PubChem3D dataset with implicit hydrogens to target distributions with explicit hydrogens. Fine-tuning EQGAT-diff for a couple of iterations further pushes state-of-the-art performance across datasets. We envision that our findings will find applications in structure-based drug design, where the accuracy of generative models for small datasets of complex molecules is critical.
翻译:深度生成扩散模型是材料科学和药物发现中从头进行3D分子设计的一条有前景的途径。然而,其效用仍受限于处理大分子结构时的次优性能以及有限的训练数据。针对这一不足,我们探索了E(3)等变扩散模型的设计空间,重点关注了此前未涉及的空白区域。通过广泛的比较分析,我们评估了连续和离散状态空间之间的相互作用。基于此项研究,我们提出了EQGAT-diff模型,该模型在QM9和GEOM-Drugs数据集上以较大优势持续超越已有模型的性能。独特之处在于,EQGAT-diff采用连续的原子位置,而化学元素和键类型为离散分类变量,并采用依赖于时间的损失加权策略,显著提升了训练收敛速度和生成样本的质量。为了进一步强化扩散模型在有限训练数据下的适用性,我们研究了在包含隐氢的大规模PubChem3D数据集上训练的EQGAT-diff模型向显氢目标分布的迁移能力。经过少数几次迭代的微调,EQGAT-diff进一步推动了各数据集上的最先进性能。我们预期,我们的研究成果将在基于结构的药物设计中得到应用,而在该领域中,针对复杂分子小数据集的生成模型的准确性至关重要。