Bayesian methods, distributionally robust optimization methods, and regularization methods are three pillars of trustworthy machine learning combating distributional uncertainty, e.g., the uncertainty of an empirical distribution compared to the true underlying distribution. This paper investigates the connections among the three frameworks and, in particular, explores why these frameworks tend to have smaller generalization errors. Specifically, first, we suggest a quantitative definition for "distributional robustness", propose the concept of "robustness measure", and formalize several philosophical concepts in distributionally robust optimization. Second, we show that Bayesian methods are distributionally robust in the probably approximately correct (PAC) sense; in addition, by constructing a Dirichlet-process-like prior in Bayesian nonparametrics, it can be proven that any regularized empirical risk minimization method is equivalent to a Bayesian method. Third, we show that generalization errors of machine learning models can be characterized using the distributional uncertainty of the nominal distribution and the robustness measures of these machine learning models, which is a new perspective to bound generalization errors, and therefore, explain the reason why distributionally robust machine learning models, Bayesian models, and regularization models tend to have smaller generalization errors in a unified manner.
翻译:贝叶斯方法、分布鲁棒优化方法和正则化方法是应对分布不确定性(例如经验分布相较于真实潜在分布的不确定性)的可信机器学习的三大支柱。本文研究这三类框架之间的联系,尤其探讨为何这些框架倾向于具有更小的泛化误差。具体而言,首先,我们提出了“分布鲁棒性”的定量定义,引入了“鲁棒性度量”的概念,并形式化了分布鲁棒优化中的若干哲理性概念。其次,我们证明贝叶斯方法在概率近似正确(PAC)意义下具有分布鲁棒性;此外,通过构造贝叶斯非参数中的狄利克雷过程先验,可以证明任何正则化的经验风险最小化方法等价于一种贝叶斯方法。第三,我们证明机器学习模型的泛化误差可利用名义分布的分布不确定性以及这些机器学习模型的鲁棒性度量来刻画,这是界定泛化误差的一种新视角,从而统一解释了为何分布鲁棒机器学习模型、贝叶斯模型和正则化模型倾向于具有更小的泛化误差。