Artificial intelligence (AI) has been widely applied in drug discovery with a major task as molecular property prediction. Despite booming techniques in molecular representation learning, fundamentals underlying molecular property prediction haven't been carefully examined yet. In this study, we conducted a systematic evaluation on a collection of representative models using various molecular representations. In addition to the commonly used MoleculeNet benchmark datasets, we also assembled a suite of opioids-related datasets from ChEMBL and two additional activity datasets from literature. To interrogate the basic predictive power, we also assembled a series of descriptors datasets with varying sizes to evaluate the models' performance. In total, we trained 62,820 models, including 50,220 models on fixed representations, 4,200 models on SMILES sequences and 8,400 models on molecular graphs. We first conducted dataset profiling and highlighted the activity-cliffs issue in the opioids-related datasets. We then conducted rigorous model evaluation and addressed key questions therein. Furthermore, we examined inter-/intra-scaffold chemical space generalization and found that activity cliffs significantly can impact prediction performance. Based on extensive experimentation and rigorous comparison, representation learning models still show limited performance in molecular property prediction in most datasets. Finally, we explored into potential causes why representation learning models fail and highlighted the importance of dataset size. By taking this respite, we reflected on the fundamentals underlying molecular property prediction, the awareness of which can, hopefully, bring better AI techniques in this field.
翻译:人工智能(AI)已广泛应用于药物发现领域,其中分子性质预测是一项核心任务。尽管分子表征学习技术蓬勃发展,但支撑分子性质预测的基础问题尚未得到细致检验。本研究对使用多种分子表征的代表性模型进行了系统评估。除常用的MoleculeNet基准数据集外,我们还从ChEMBL数据库中整理了一套阿片类药物相关数据集,并从文献中获取了另外两个活性数据集。为探究基本预测能力,我们进一步构建了一系列不同规模的描述符数据集以评估模型性能。总计训练了62,820个模型,包括50,220个基于固定表征的模型、4,200个基于SMILES序列的模型和8,400个基于分子图的模型。我们首先进行数据集剖析,重点揭示了阿片类药物相关数据集中存在的活性悬崖问题。随后开展严格的模型评估,解答了其中的关键问题。此外,我们考察了跨骨架/骨架内化学空间泛化能力,发现活性悬崖会显著影响预测性能。基于大量实验和严谨比较发现,在大多数数据集中,表征学习模型在分子性质预测上仍表现有限。最后,我们探讨了表征学习模型失效的潜在原因,并强调数据集规模的重要性。通过此次反思,我们重新审视了分子性质预测的基础问题,希望这些认知能为该领域带来更优的AI技术。