Machine learning is at the heart of managing the real-world problems associated with massive data. With the success of neural networks on such large-scale problems, more research in machine learning is being conducted now than ever before. This dissertation focuses on three different projects rooted in mathematical theory for machine learning applications. The first project deals with supervised learning and manifold learning. In theory, one of the main problems in supervised learning is that of function approximation: that is, given some data set $\mathcal{D}=\{(x_j,f(x_j))\}_{j=1}^M$, can one build a model $F\approx f$? We introduce a method which aims to remedy several of the theoretical shortcomings of the current paradigm for supervised learning. The second project deals with transfer learning, which is the study of how an approximation process or model learned on one domain can be leveraged to improve the approximation on another domain. We study such liftings of functions when the data is assumed to be known only on a part of the whole domain. We are interested in determining subsets of the target data space on which the lifting can be defined, and how the local smoothness of the function and its lifting are related. The third project is concerned with the classification task in machine learning, particularly in the active learning paradigm. Classification has often been treated as an approximation problem as well, but we propose an alternative approach leveraging techniques originally introduced for signal separation problems. We introduce theory to unify signal separation with classification and a new algorithm which yields competitive accuracy to other recent active learning algorithms while providing results much faster.
翻译:机器学习是处理海量数据相关现实问题的核心。随着神经网络在此类大规模问题上的成功,当前机器学习领域的研究比以往任何时候都更加活跃。本论文聚焦于三个植根于机器学习应用的数学理论的不同项目。第一个项目涉及监督学习与流形学习。理论上,监督学习的主要问题之一是函数逼近问题:即给定数据集 $\mathcal{D}=\{(x_j,f(x_j))\}_{j=1}^M$,能否构建模型 $F\approx f$?我们提出一种方法,旨在修正当前监督学习范式的若干理论缺陷。第二个项目研究迁移学习,即探索如何利用在一个领域学习的逼近过程或模型来提升另一领域的逼近效果。我们研究当数据仅在整个域的部分区域已知时函数的提升问题,重点确定提升可定义的目标数据空间子集,并分析函数与其提升的局部光滑性之间的关联。第三个项目关注机器学习中的分类任务,特别是在主动学习范式下的分类。分类问题常被视为逼近问题,但我们提出一种替代方案,利用最初为信号分离问题开发的技术。我们建立了统一信号分离与分类的理论框架,并提出一种新算法,该算法在保持与近期其他主动学习算法相当精度的同时,能显著提升计算效率。