New powerful tools for tackling life science problems have been created by recent advances in machine learning. The purpose of the paper is to discuss the potential advantages of gene recommendation performed by artificial intelligence (AI). Indeed, gene recommendation engines try to solve this problem: if the user is interested in a set of genes, which other genes are likely to be related to the starting set and should be investigated? This task was solved with a custom deep learning recommendation engine, DeepProphet2 (DP2), which is freely available to researchers worldwide via https://www.generecommender.com?utm_source=DeepProphet2_paper&utm_medium=pdf. Hereafter, insights behind the algorithm and its practical applications are illustrated. The gene recommendation problem can be addressed by mapping the genes to a metric space where a distance can be defined to represent the real semantic distance between them. To achieve this objective a transformer-based model has been trained on a well-curated freely available paper corpus, PubMed. The paper describes multiple optimization procedures that were employed to obtain the best bias-variance trade-off, focusing on embedding size and network depth. In this context, the model's ability to discover sets of genes implicated in diseases and pathways was assessed through cross-validation. A simple assumption guided the procedure: the network had no direct knowledge of pathways and diseases but learned genes' similarities and the interactions among them. Moreover, to further investigate the space where the neural network represents genes, the dimensionality of the embedding was reduced, and the results were projected onto a human-comprehensible space. In conclusion, a set of use cases illustrates the algorithm's potential applications in a real word setting.
翻译:机器学习的最新进展为生命科学问题的研究创造了强大的新工具。本文旨在探讨人工智能(AI)进行基因推荐的潜在优势。具体而言,基因推荐引擎试图解决以下问题:若用户对一组基因感兴趣,哪些其他基因可能与起始集合相关并值得进一步研究?该任务通过定制化的深度学习推荐引擎DeepProphet2(DP2)实现,该引擎现通过https://www.generecommender.com?utm_source=DeepProphet2_paper&utm_medium=pdf免费向全球研究人员开放。下文将阐述该算法背后的原理及其实际应用。基因推荐问题可通过将基因映射到一个度量空间来解决,在该空间中可定义距离以表示基因之间的真实语义距离。为实现这一目标,我们基于Transformer模型,在精心筛选的免费公开论文语料库PubMed上进行了训练。本文描述了为获得最优偏置-方差权衡而采用的多种优化策略,重点关注嵌入大小和网络深度。在此背景下,通过交叉验证评估了模型发现与疾病和通路相关的基因集合的能力。整个过程遵循一个简单假设:网络对通路和疾病没有直接知识,而是通过学习基因之间的相似性及其相互作用来运作。此外,为深入探究神经网络表示基因的空间,我们降低了嵌入的维度,并将结果投影至人类可理解的空间中。最后,通过一系列用例展示了该算法在实际场景中的潜在应用。