We demonstrate that the forecasting combination puzzle is a consequence of the methodology commonly used to produce forecast combinations. By the combination puzzle, we refer to the empirical finding that predictions formed by combining multiple forecasts in ways that seek to optimize forecast performance often do not out-perform more naive, e.g. equally-weighted, approaches. In particular, we demonstrate that, due to the manner in which such forecasts are typically produced, tests that aim to discriminate between the predictive accuracy of competing combination strategies can have low power, and can lack size control, leading to an outcome that favours the naive approach. We show that this poor performance is due to the behavior of the corresponding test statistic, which has a non-standard asymptotic distribution under the null hypothesis of no inferior predictive accuracy, rather than the {standard normal distribution that is} {typically adopted}. In addition, we demonstrate that the low power of such predictive accuracy tests in the forecast combination setting can be completely avoided if more efficient estimation strategies are used in the production of the combinations, when feasible. We illustrate these findings both in the context of forecasting a functional of interest and in terms of predictive densities. A short empirical example {using daily financial returns} exemplifies how researchers can avoid the puzzle in practical settings.
翻译:我们证明,预测组合之谜是通常用于生成预测组合的方法论所导致的结果。所谓“组合之谜”,指的是一个经验性发现:通过优化预测性能的方式组合多个预测所形成的预测结果,往往并不优于更朴素的方法(例如等权重方法)。具体而言,我们证明,由于此类预测的产生方式,旨在区分竞争性组合策略预测准确性的检验可能具有低检验功效,且可能缺乏尺度控制,从而导向偏好朴素方法的结果。我们表明,这种低劣表现源于相应检验统计量的行为:在“无劣化预测准确性”的原假设下,该统计量服从非标准渐近分布,而非通常采用的{标准正态分布}。此外,我们证明,若在生成组合时(在可行情况下)采用更高效的估计策略,则可完全避免预测组合设定中此类预测准确性检验的低功效问题。我们通过预测感兴趣的函数以及预测密度两种情境,阐明这些发现。一个关于{日度金融收益率}的简短实证示例,展示了研究者如何在实践场景中规避该谜题。