To incorporate useful information from related statistical tasks into the target one, we propose a two-step transfer learning algorithm in the general M-estimators framework with decomposable regularizers in high-dimensions. When the informative sources are known in the oracle sense, in the first step, we acquire knowledge of the target parameter by pooling the useful source datasets and the target one. In the second step, the primal estimator is fine-tuned using the target dataset. In contrast to the existing literatures which exert homogeneity conditions for Hessian matrices of various population loss functions, our theoretical analysis shows that even if the Hessian matrices are heterogeneous, the pooling estimators still provide adequate information by slightly enlarging regularization, and numerical studies further validate the assertion. Sparse regression and low-rank trace regression for both linear and generalized linear cases are discussed as two specific examples under the M-estimators framework. When the informative source datasets are unknown, a novel truncated-penalized algorithm is proposed to acquire the primal estimator and its oracle property is proved. Extensive numerical experiments are conducted to support the theoretical arguments.
翻译:为将相关统计任务中的有用信息纳入目标任务,我们提出了一种面向高维中具有可分解正则项的一般M-估计量框架下的两步迁移学习算法。当已知有效源数据集时,第一步通过合并有用源数据集与目标数据集获取目标参数的初步知识;第二步利用目标数据集对初步估计量进行微调。与现有文献中对不同总体损失函数的Hessian矩阵施加同质性条件不同,我们的理论分析表明,即使Hessian矩阵存在异质性,通过略微扩大正则化参数,合并估计量仍能提供充分信息,数值研究进一步验证了这一论断。在M-估计量框架下,我们以稀疏回归以及线性和广义线性情况下的低秩迹回归作为两个具体例子展开讨论。当有效源数据集未知时,我们提出了一种新颖的截断惩罚算法来获取初步估计量,并证明了其Oracle性质。大量数值实验支持了理论论证。