The empirical risk minimization (ERM) problem with relative entropy regularization (ERM-RER) is investigated under the assumption that the reference measure is a {\sigma}-finite measure, and not necessarily a probability measure. Under this assumption, which leads to a generalization of the ERM-RER problem allowing a larger degree of flexibility for incorporating prior knowledge, numerous relevant properties are stated. Among these properties, the solution to this problem, if it exists, is shown to be a unique probability measure, often mutually absolutely continuous with the reference measure. Such a solution exhibits a probably-approximately-correct guarantee for the ERM problem independently of whether the latter possesses a solution. For a fixed dataset, the empirical risk is shown to be a sub-Gaussian random variable when the models are sampled from the solution to the ERM-RER problem. The generalization capabilities of the solution to the ERM-RER problem (the Gibbs algorithm) are studied via the sensitivity of the expected empirical risk to deviations from such a solution towards alternative probability measures. Finally, an interesting connection between sensitivity, generalization error, and lautum information is established
翻译:本文研究了具有相对熵正则化(ERM-RER)的经验风险最小化问题,其中参考测度假定为σ-有限测度,而非必然是概率测度。在此假设下,ERM-RER问题得到推广,从而允许在融入先验知识时具备更高的灵活性,并阐明了若干重要性质。其中,该问题的解(若存在)被证明为唯一概率测度,且通常与参考测度相互绝对连续。该解独立于原经验风险最小化问题是否有解,均能提供一种“可能近似正确”(PAC)保证。对于固定数据集,当模型从ERM-RER问题的解中采样时,经验风险被证明为次高斯随机变量。通过分析期望经验风险对偏离该解(朝向替代概率测度)的敏感性,研究了ERM-RER问题解(即吉布斯算法)的泛化能力。最后,建立了敏感性、泛化误差与劳图姆信息之间的有趣联系。