Predictions from machine learning algorithms can vary across random seeds, inducing instability in downstream debiased machine learning estimators. We formalize random seed stability via a concentration condition and prove that subbagging guarantees stability for any bounded-outcome regression algorithm. We introduce a new cross-fitting procedure, adaptive cross-bagging, which simultaneously eliminates seed dependence from both nuisance estimation and sample splitting in debiased machine learning. Numerical experiments confirm that the method achieves the targeted level of stability whereas alternatives do not. Our method incurs a small computational penalty relative to standard practice whereas alternative methods incur large penalties.
翻译:机器学习算法的预测结果可能因随机种子不同而产生差异,进而在下游去偏机器学习估计器中引发不稳定性。我们通过浓度条件形式化随机种子的稳定性,并证明子袋装法能保证任意有界输出回归算法的稳定性。我们引入一种新的交叉拟合过程——自适应交叉袋装法,该方法能同时消除去偏机器学习中因干扰估计和样本分割引起的种子依赖性。数值实验证实,该方法能够达到目标稳定性水平,而替代方法则无法实现。与标准做法相比,我们的方法仅增加少量计算成本,而替代方法则需承担高昂的计算代价。