The success of drug discovery and development relies on the precise prediction of molecular activities and properties. While in silico molecular property prediction has shown remarkable potential, its use has been limited so far to assays for which large amounts of data are available. In this study, we use a fine-tuned large language model to integrate biological assays based on their textual information, coupled with Barlow Twins, a Siamese neural network using a novel self-supervised learning approach. This architecture uses both assay information and molecular fingerprints to extract the true molecular information. TwinBooster enables the prediction of properties of unseen bioassays and molecules by providing state-of-the-art zero-shot learning tasks. Remarkably, our artificial intelligence pipeline shows excellent performance on the FS-Mol benchmark. This breakthrough demonstrates the application of deep learning to critical property prediction tasks where data is typically scarce. By accelerating the early identification of active molecules in drug discovery and development, this method has the potential to help streamline the identification of novel therapeutics.
翻译:药物发现与开发的成功依赖于对分子活性与性质的精确预测。尽管计算机模拟分子性质预测已展现出显著潜力,但其应用迄今仍局限于具备大量数据的测定实验。本研究中,我们利用微调后的大语言模型,基于文本信息整合生物测定实验,并结合采用新型自监督学习方法的孪生神经网络Barlow Twins。该架构同时利用测定实验信息与分子指纹提取真正的分子信息。TwinBooster通过提供最先进的零样本学习任务,能够预测未见生物测定实验与分子的性质。值得注意的是,我们的人工智能流水线在FS-Mol基准测试中展现出卓越性能。这一突破证明了深度学习在数据通常稀缺的关键性质预测任务中的适用性。通过加速药物发现与开发中活性分子的早期识别,该方法有望助力新型治疗药物的高效筛选。