Trustworthy machine learning aims at combating distributional uncertainties in training data distributions compared to population distributions. Typical treatment frameworks include the Bayesian approach, (min-max) distributionally robust optimization (DRO), and regularization. However, two issues have to be raised: 1) All these methods are biased estimators of the true optimal cost; 2) the prior distribution in the Bayesian method, the radius of the distributional ball in the DRO method, and the regularizer in the regularization method are difficult to specify. This paper studies a new framework that unifies the three approaches and that addresses the two challenges mentioned above. The asymptotic properties (e.g., consistency and asymptotic normalities), non-asymptotic properties (e.g., unbiasedness and generalization error bound), and a Monte--Carlo-based solution method of the proposed model are studied. The new model reveals the trade-off between the robustness to the unseen data and the specificity to the training data.
翻译:可信机器学习旨在对抗训练数据分布与总体分布之间的分布不确定性。典型的处理框架包括贝叶斯方法、(最小-最大)分布鲁棒优化(DRO)以及正则化方法。然而,需要指出两个问题:1)所有这些方法都是真实最优成本的有偏估计;2)贝叶斯方法中的先验分布、DRO方法中的分布球半径以及正则化方法中的正则化因子难以确定。本文研究了一个统一上述三种方法的新框架,并解决了上述两个挑战。本文研究了所提模型的渐近性质(如一一致性和渐近正态性)、非渐近性质(如无偏性和泛化误差界)以及基于蒙特卡洛的求解方法。新模型揭示了对于未见数据的鲁棒性与对训练数据的特异性之间的权衡关系。