Recently deep learning based quantitative structure-activity relationship (QSAR) models has shown surpassing performance than traditional methods for property prediction tasks in drug discovery. However, most DL based QSAR models are restricted to limited labeled data to achieve better performance, and also are sensitive to model scale and hyper-parameters. In this paper, we propose Uni-QSAR, a powerful Auto-ML tool for molecule property prediction tasks. Uni-QSAR combines molecular representation learning (MRL) of 1D sequential tokens, 2D topology graphs, and 3D conformers with pretraining models to leverage rich representation from large-scale unlabeled data. Without any manual fine-tuning or model selection, Uni-QSAR outperforms SOTA in 21/22 tasks of the Therapeutic Data Commons (TDC) benchmark under designed parallel workflow, with an average performance improvement of 6.09\%. Furthermore, we demonstrate the practical usefulness of Uni-QSAR in drug discovery domains.
翻译:近年来,基于深度学习的定量构效关系(QSAR)模型在药物发现领域的性质预测任务中展现出超越传统方法的性能。然而,大多数基于深度学习的QSAR模型受限于有限的标注数据,难以取得更优性能,且对模型规模和超参数高度敏感。本文提出Uni-QSAR——一种面向分子性质预测任务的强大自动机器学习工具。Uni-QSAR整合了一维序列令牌、二维拓扑图以及三维构象的分子表征学习(MRL),并融合预训练模型以利用大规模无标注数据的丰富表征。无需任何手动微调或模型选择,Uni-QSAR在治疗数据共享平台(TDC)基准测试的22项任务中于21项上超越当前最优方法(SOTA),在设计的并行工作流程下平均性能提升6.09%。此外,我们验证了Uni-QSAR在药物发现领域的实际应用价值。