Outliers widely occur in big-data applications and may severely affect statistical estimation and inference. In this paper, a framework of outlier-resistant estimation is introduced to robustify an arbitrarily given loss function. It has a close connection to the method of trimming and includes explicit outlyingness parameters for all samples, which in turn facilitates computation, theory, and parameter tuning. To tackle the issues of nonconvexity and nonsmoothness, we develop scalable algorithms with implementation ease and guaranteed fast convergence. In particular, a new technique is proposed to alleviate the requirement on the starting point such that on regular datasets, the number of data resamplings can be substantially reduced. Based on combined statistical and computational treatments, we are able to perform nonasymptotic analysis beyond M-estimation. The obtained resistant estimators, though not necessarily globally or even locally optimal, enjoy minimax rate optimality in both low dimensions and high dimensions. Experiments in regression, classification, and neural networks show excellent performance of the proposed methodology at the occurrence of gross outliers.
翻译:离群值在大数据应用中普遍存在,可能严重影响统计估计与推断。本文提出一个离群鲁棒性估计框架,可对任意给定的损失函数进行鲁棒化处理。该框架与修剪法紧密相关,并为所有样本引入显式离群性参数,从而便于计算、理论分析和参数调优。针对非凸性和非光滑性问题,我们开发了易于实现且保证快速收敛的可扩展算法。特别地,我们提出一种新技术来降低对起始点的要求,从而在常规数据集上大幅减少数据重采样次数。通过统计与计算的联合处理,我们得以进行超越M估计的非渐近分析。所获得的鲁棒估计量虽然未必达到全局甚至局部最优,但在低维和高维场景中均能实现极小极大速率最优性。在回归、分类和神经网络上的实验表明,所提方法在出现严重离群值时具有卓越性能。