Sample reweighting is one of the most widely used methods for correcting the error of least squares learning algorithms in reproducing kernel Hilbert spaces (RKHS), that is caused by future data distributions that are different from the training data distribution. In practical situations, the sample weights are determined by values of the estimated Radon-Nikod\'ym derivative, of the future data distribution w.r.t.~the training data distribution. In this work, we review known error bounds for reweighted kernel regression in RKHS and obtain, by combination, novel results. We show under weak smoothness conditions, that the amount of samples, needed to achieve the same order of accuracy as in the standard supervised learning without differences in data distributions, is smaller than proven by state-of-the-art analyses.
翻译:样本重加权是纠正再生核希尔伯特空间(RKHS)中最小二乘学习算法误差最广泛使用的方法之一,这些误差源于未来数据分布与训练数据分布的不同。在实际情况下,样本权重由未来数据分布相对于训练数据分布的估计Radon-Nikodým导数的值决定。本文回顾了RKHS中重加权核回归的已知误差界,并通过组合获得了新的结果。我们证明了在弱光滑性条件下,达到与标准监督学习(无数据分布差异)相同精度所需的样本量,比现有最优分析所证明的要少。