Obtaining guarantees on the convergence of the minimizers of empirical risks to the ones of the true risk is a fundamental matter in statistical learning. Instead of deriving guarantees on the usual estimation error, the goal of this paper is to provide concentration inequalities on the distance between the sets of minimizers of the risks for a broad spectrum of estimation problems. In particular, the risks are defined on metric spaces through probability measures that are also supported on metric spaces. A particular attention will therefore be given to include unbounded spaces and non-convex cost functions that might also be unbounded. This work identifies a set of assumptions allowing to describe a regime that seem to govern the concentration in many estimation problems, where the empirical minimizers are stable. This stability can then be leveraged to prove parametric concentration rates in probability and in expectation. The assumptions are verified, and the bounds showcased, on a selection of estimation problems such as barycenters on metric space with positive or negative curvature, subspaces of covariance matrices, regression problems and entropic-Wasserstein barycenters.
翻译:为经验风险的最小化器收敛到真实风险的最小化器提供保证是统计学习中的一个基本问题。本文的目标不是推导通常的估计误差保证,而是为一系列广泛的估计问题中风险最小化器集合之间的距离提供集中不等式。特别地,风险定义在度量空间上,并通过同样支撑在度量空间上的概率测度来定义。因此,本文将特别关注包含无界空间以及可能也无界的非凸代价函数。这项工作识别了一组假设,这些假设允许描述一种在许多估计问题中似乎支配着集中性的状态,其中经验最小化器是稳定的。然后可以利用这种稳定性来证明概率和期望上的参数化集中速度。这些假设得到了验证,并且边界在选定的估计问题上进行了展示,例如正曲率或负曲率度量空间中的重心、协方差矩阵的子空间、回归问题以及熵-Wasserstein重心。