While sequential residual fitting is the bedrock of standard boosting frameworks, it inherently breeds learner redundancy by repeatedly revisiting correlated error components. To address this bottleneck, we propose a shift from residual fitting to \textit{residual orthogonalization} and introduce SCBoost. Our framework tackles redundancy through two complementary mechanisms: Spectral Residual Projection (SRP) and Covariance-Regularized Weighting (CRW). During training, SRP projects each residual target onto the orthogonal complement of the historical prediction subspace, forcing successive learners to capture only novel empirical innovations. During aggregation, CRW optimizes ensemble weights on a validation set with an explicit covariance penalty to mitigate remaining correlations. Theoretically, we provide a finite-sample geometric characterization proving that SRP yields an exact additive residual-energy decomposition. Furthermore, under an isotropic-noise assumption, we rigorously establish the conditions under which this projection improves the effective Signal-to-Noise Ratio. Extensive experiments across ten benchmark datasets demonstrate that SCBoost delivers strong out-of-the-box performance, particularly in accuracy and F1 score. This work reinterprets boosting through a geometric lens, suggesting that explicit redundancy control is a principled and necessary step toward more efficient ensemble architectures.
翻译:虽然序贯残差拟合是标准提升框架的基石,但该方法会因反复处理相关误差分量而内生地滋生学习器冗余性。为解决这一瓶颈,我们提出从残差拟合转向*残差正交化*,并引入SCBoost框架。该框架通过两种互补机制消除冗余:谱残差投影(SRP)和协方差正则化加权(CRW)。训练阶段,SRP将每个残差目标投影到历史预测子空间的正交补空间上,迫使后续学习器仅捕捉新的经验创新成分;聚合阶段,CRW在验证集上通过显式协方差惩罚项优化集成权重,以减轻剩余相关性。理论上,我们提供了有限样本几何特性刻画,证明SRP可实现精确的加性残差能量分解。进一步地,在各向同性噪声假设下,我们严格建立了该投影能提升有效信噪比的条件。在十个基准数据集上的大量实验表明,SCBoost展现出优异的开箱即用性能,尤其在准确率和F1分数方面表现突出。本工作通过几何视角重新诠释了提升方法,揭示显式冗余控制是迈向更高效集成架构的原则性且必要的步骤。