In recent years, pretraining models have made significant advancements in the fields of natural language processing (NLP), computer vision (CV), and life sciences. The significant advancements in NLP and CV are predominantly driven by the expansion of model parameters and data size, a phenomenon now recognized as the scaling laws. However, research exploring scaling law in molecular pretraining models remains unexplored. In this work, we present Uni-Mol2 , an innovative molecular pretraining model that leverages a two-track transformer to effectively integrate features at the atomic level, graph level, and geometry structure level. Along with this, we systematically investigate the scaling law within molecular pretraining models, characterizing the power-law correlations between validation loss and model size, dataset size, and computational resources. Consequently, we successfully scale Uni-Mol2 to 1.1 billion parameters through pretraining on 800 million conformations, making it the largest molecular pretraining model to date. Extensive experiments show consistent improvement in the downstream tasks as the model size grows. The Uni-Mol2 with 1.1B parameters also outperforms existing methods, achieving an average 27% improvement on the QM9 and 14% on COMPAS-1D dataset.
翻译:近年来,预训练模型在自然语言处理(NLP)、计算机视觉(CV)和生命科学领域取得了显著进展。NLP和CV领域的重大进步主要由模型参数和数据规模的扩展所驱动,这一现象现在被公认为缩放定律。然而,探索分子预训练模型中缩放定律的研究仍然处于空白。在本工作中,我们提出了Uni-Mol2,这是一个创新的分子预训练模型,它利用双轨Transformer有效地整合了原子级别、图级别和几何结构级别的特征。与此同时,我们系统地研究了分子预训练模型内部的缩放定律,刻画了验证损失与模型大小、数据集大小和计算资源之间的幂律相关性。因此,我们通过在8亿个构象上进行预训练,成功地将Uni-Mol2扩展到11亿参数,使其成为迄今为止最大的分子预训练模型。大量实验表明,随着模型规模的增长,下游任务的性能得到了一致的提升。拥有11亿参数的Uni-Mol2也超越了现有方法,在QM9数据集上平均提升了27%,在COMPAS-1D数据集上平均提升了14%。