Finding new drugs is getting harder and harder. One of the hopes of drug discovery is to use machine learning models to predict molecular properties. That is why models for molecular property prediction are being developed and tested on benchmarks such as MoleculeNet. However, existing benchmarks are unrealistic and are too different from applying the models in practice. We have created a new practical \emph{Lo-Hi} benchmark consisting of two tasks: Lead Optimization (Lo) and Hit Identification (Hi), corresponding to the real drug discovery process. For the Hi task, we designed a novel molecular splitting algorithm that solves the Balanced Vertex Minimum $k$-Cut problem. We tested state-of-the-art and classic ML models, revealing which works better under practical settings. We analyzed modern benchmarks and showed that they are unrealistic and overoptimistic. Review: https://openreview.net/forum?id=H2Yb28qGLV Lo-Hi benchmark: https://github.com/SteshinSS/lohi_neurips2023 Lo-Hi splitter library: https://github.com/SteshinSS/lohi_splitter
翻译:发现新药正变得越来越困难。药物发现的希望之一是利用机器学习模型预测分子性质。因此,分子性质预测模型在MoleculeNet等基准上进行开发和测试。然而,现有基准不切实际,与模型在实践中的应用相去甚远。我们创建了一个新的实用**Lo-Hi**基准,包含两个任务:先导化合物优化(Lead Optimization, Lo)和命中化合物识别(Hit Identification, Hi),对应真实的药物发现过程。对于Hi任务,我们设计了一种新颖的分子分割算法,解决了平衡顶点最小$k$-切割问题。我们测试了最先进和经典机器学习模型,揭示了哪些模型在实际环境中表现更优。我们分析了现代基准,证明它们不切实际且过于乐观。评审:https://openreview.net/forum?id=H2Yb28qGLV Lo-Hi基准:https://github.com/SteshinSS/lohi_neurips2023 Lo-Hi分割器库:https://github.com/SteshinSS/lohi_splitter