We introduce Prior Knowledge Acceleration (PKA), a batch-update method for variance that reuses previously computed sufficient statistics to avoid full recomputation. The update identity is algebraically equivalent to the pairwise formula of Chan, Golub, and LeVeque (1983); our contribution is a runtime-cost analysis that derives an explicit acceleration factor $τ_a$ and identifies the data-size regime where batch updating outperforms both naïve recomputation and Ross's single-sample method. We prove that Ross's approach is preferable only when the new batch contains a single sample ($N_2 = 1$). We further generalise the framework to covariance and other decomposable statistics. Benchmarks against Welford, Chan pairwise, and naïve two-pass baselines on synthetic and real-world streaming data confirm the theoretical predictions, with speedups of up to $454\times$ when the prior dataset is large relative to the new batch.
翻译:我们提出先验知识加速法(Prior Knowledge Acceleration, PKA),这是一种面向方差的批量更新方法,通过复用先前计算的充分统计量来避免完全重计算。该更新恒等式在代数上等价于Chan、Golub与LeVeque(1983)提出的成对公式;我们的贡献在于运行时成本分析,由此推导出显式加速因子$τ_a$,并识别出批量更新优于朴素重计算及Ross单样本方法的数据规模区间。我们证明,仅当新批数据包含单一样本($N_2 = 1$)时,Ross方法才更具优势。进一步地,我们将该框架推广至协方差及其他可分解统计量。在合成数据与真实流式数据上的基准测试中,该方法相较于Welford算法、Chan成对算法及朴素双通基线符合理论预测:当先验数据集相对新批数据规模较大时,加速比可达$454\times$。