We study the fundamental problem of calibrating a linear binary classifier of the form $σ(\hat{w}^\top x)$, where the feature vector $x$ is Gaussian, $σ$ is a link function, and $\hat{w}$ is an estimator of the true linear weight $w^\star$. By interpolating with a noninformative $\textit{chance classifier}$, we construct a well-calibrated predictor whose interpolation weight depends on the angle $\angle(\hat{w}, w_\star)$ between the estimator $\hat{w}$ and the true linear weight $w_\star$. We establish that this angular calibration approach is provably well-calibrated in a high-dimensional regime where the number of samples and features both diverge, at a comparable rate. The angle $\angle(\hat{w}, w_\star)$ can be consistently estimated. Furthermore, the resulting predictor is uniquely $\textit{Bregman-optimal}$, minimizing the Bregman divergence to the true label distribution within a suitable class of calibrated predictors. Our work is the first to provide a calibration strategy that satisfies both calibration and optimality properties provably in high dimensions. Additionally, we identify conditions under which a classical Platt-scaling predictor converges to our Bregman-optimal calibrated solution. Thus, Platt-scaling also inherits these desirable properties provably in high dimensions.
翻译:我们研究形如 $σ(\hat{w}^\top x)$ 的线性二分类器校准的基本问题,其中特征向量 $x$ 服从高斯分布,$σ$ 为链接函数,$\hat{w}$ 是真值线性权重 $w^\star$ 的估计量。通过与非信息性的 **机会分类器** 进行插值,我们构建了一个良好校准的预测器,其插值权重取决于估计量 $\hat{w}$ 与真值线性权重 $w_\star$ 之间的角度 $\angle(\hat{w}, w_\star)$。我们证明该角度校准方法在高维场景下是可证明的良好校准——其中样本量与特征量以可比速率同时发散。角度 $\angle(\hat{w}, w_\star)$ 可被一致地估计。此外,该预测器具有唯一的 **Bregman最优性**,即在某类合适的校准预测器中,它使预测分布与真实标签分布之间的Bregman散度最小化。本研究首次提出了一种能在高维条件下同时满足校准性与最优性(且具有可证明性)的校准策略。同时,我们识别了经典Platt缩放预测器收敛至该Bregman最优校准解的条件。因此,Platt缩放在高维场景下同样继承了这些理想性质(且具有可证明性)。