Applying machine learning to biological sequences - DNA, RNA and protein - has enormous potential to advance human health, environmental sustainability, and fundamental biological understanding. However, many existing machine learning methods are ineffective or unreliable in this problem domain. We study these challenges theoretically, through the lens of kernels. Methods based on kernels are ubiquitous: they are used to predict molecular phenotypes, design novel proteins, compare sequence distributions, and more. Many methods that do not use kernels explicitly still rely on them implicitly, including a wide variety of both deep learning and physics-based techniques. While kernels for other types of data are well-studied theoretically, the structure of biological sequence space (discrete, variable length sequences), as well as biological notions of sequence similarity, present unique mathematical challenges. We formally analyze how well kernels for biological sequences can approximate arbitrary functions on sequence space and how well they can distinguish different sequence distributions. In particular, we establish conditions under which biological sequence kernels are universal, characteristic and metrize the space of distributions. We show that a large number of existing kernel-based machine learning methods for biological sequences fail to meet our conditions and can as a consequence fail severely. We develop straightforward and computationally tractable ways of modifying existing kernels to satisfy our conditions, imbuing them with strong guarantees on accuracy and reliability. Our proof techniques build on and extend the theory of kernels with discrete masses. We illustrate our theoretical results in simulation and on real biological data sets.
翻译:将机器学习应用于生物序列(DNA、RNA和蛋白质)具有极大潜力,可推动人类健康、环境可持续性及基础生物学认知的发展。然而,许多现有机器学习方法在该问题域中表现低效或不可靠。我们通过核函数的视角,从理论上研究这些挑战。基于核函数的方法应用广泛:它们被用于预测分子表型、设计新型蛋白质、比较序列分布等。许多未显式使用核函数的方法仍隐含依赖于它们,包括各类深度学习及基于物理学的技术。尽管其他数据类型对应的核函数在理论上已得到充分研究,但生物序列空间的结构(离散、变长序列)以及生物学序列相似性概念提出了独特的数学挑战。我们系统分析了生物序列核函数在序列空间上逼近任意函数的能力,以及区分不同序列分布的效果。具体而言,我们建立了生物序列核函数具备普适性、特征性并能度量分布空间的充分条件。研究表明,现有大量基于核函数的生物序列机器学习方法未能满足我们的条件,因此可能产生严重失效。我们开发了直接且计算可行的途径来修改现有核函数,使其满足条件,从而赋予其准确性与可靠性的强保证。我们的证明方法基于并扩展了离散质量核函数的理论。在仿真实验及真实生物数据集上,我们验证了理论结果。