The main goal in regression modelling consists in approximating the conditional mean of a response given a set of features. A regression function is said to be calibrated if the resulting mean estimates match the true conditional means for almost every set of features. Aiming for calibration seems not achievable in practice as one typically deals with finite samples of noisy observations. A weaker notion of calibration is auto-calibration, and it means that the expectation of responses being given the same mean estimate matches this estimate. This notion is important, e.g., in insurance pricing as it ensures no cross-subsidization between different price cohorts. In this paper, we show that boosting trees can be used to test necessary conditions for calibration and auto-calibration, respectively. The practical relevance of our approach is supported by a numerical example, in which the proposed tests prove to be very powerful on a large insurance dataset.
翻译:回归建模的主要目标是近似给定一组特征时响应的条件均值。若回归函数产生的均值估计与几乎所有特征集上的真实条件均值相匹配,则称该回归函数已校准。在实践中,由于通常处理的是有限样本的含噪观测,追求校准似乎难以实现。一种较弱的校准概念是自校准,它意味着具有相同均值估计的响应的期望与该估计值相匹配。这一概念在保险定价等领域尤为重要,因为它确保了不同定价群组之间不存在交叉补贴。在本文中,我们展示了提升树可分别用于检验校准与自校准的必要条件。我们方法的实际相关性通过一个数值示例得到支持,该示例显示,所提出的检验方法在一个大型保险数据集上具有非常高的检验效力。