Model-assisted regression estimation is fundamental in survey sampling for incorporating auxiliary information. However, when the auxiliary dimension grows with the sample size, the standard Generalized regression (GREG) estimator can exhibit non-negligible bias under informative sampling, even when the working model is correctly specified. This failure stems from the double use of sampled outcomes simultaneously for fitting the regression and for forming the residual correction. We propose a sample-split REGression (SREG) estimator based on K-fold cross-fitting that eliminates this bias by pairing each unit's residual with an out-of-fold prediction. The resulting estimator is first-order equivalent to the oracle difference estimator under a weak prediction-norm consistency requirement, without requiring root-n consistent estimation of regression coefficients. We establish asymptotic normality and prove consistency of a variance estimator based on cross-fitted residuals. The key conditional fluctuation assumption is verified for simple random, stratified, and rejective sampling. Simulations demonstrate that SREG effectively removes high-dimensional bias while maintaining competitive efficiency.
翻译:模型辅助回归估计是调查抽样中利用辅助信息的基础方法。然而,当辅助变量维度随样本量增长时,即使在指定正确的工作模型下,标准广义回归估计量在信息性抽样中仍可能产生不可忽视的偏差。这一缺陷源于回归拟合与残差修正过程中对样本结果的重复使用。本文提出基于K折交叉拟合的样本分割回归估计量,通过将每个单元的残差与折外预测配对来消除该偏差。在弱预测范数一致性条件下,该估计量在渐近等价性与基准差异估计量一阶等价,且无需回归系数的根号n相合估计。我们建立了渐近正态性,并证明了基于交叉拟合残差的方差估计量的一致性。验证了简单随机抽样、分层抽样与拒绝抽样下的关键条件波动假设。模拟实验表明,SREG在保持竞争性效率的同时有效消除了高维偏差。