Forecasts crave a rating that reflects the forecast's quality in the context of what is possible in theory and what is reasonable to expect in practice. Granular forecasts in the regime of low count rates - as they often occur in retail, for which an intermittent demand of a handful might be observed per product, day, and location - are dominated by the inevitable statistical uncertainty of the Poisson distribution. This makes it hard to judge whether a certain metric value is dominated by Poisson noise or truly indicates a bad prediction model. To make things worse, every evaluation metric suffers from scaling: Its value is mostly defined by the predicted selling rate and the resulting rate-dependent Poisson noise, and only secondarily by the quality of the forecast. For any metric, comparing two groups of forecasted products often yields "the slow movers are performing worse than the fast movers" or vice versa - the na\"ive scaling trap. To distill the intrinsic quality of a forecast, we stratify predictions into buckets of approximately equal rate and evaluate metrics for each bucket separately. By comparing the achieved value per bucket to benchmarks, we obtain a scaling-aware rating of count forecasts. Our procedure avoids the na\"ive scaling trap, provides an immediate intuitive judgment of forecast quality, and allows to compare forecasts for different products or even industries.
翻译:预测渴望一种评级,这种评级能反映预测质量,既要考虑理论上的可能性,也要考虑实践中的合理预期。低计数率下的细粒度预测——例如零售中常见的情况,即每个产品、每天、每个地点可能只观察到零星的非连续需求——受到泊松分布固有统计不确定性的主导。这使得难以判断某个特定指标值是否主要由泊松噪声主导,还是真正表明预测模型不佳。更糟糕的是,每种评估指标都受缩放效应影响:其数值主要由预测销售率和由此产生的率相关泊松噪声决定,而预测质量只占次要因素。对于任何指标,比较两组预测产品通常会产生“慢销产品表现比快销产品差”或相反的结果——这就是朴素缩放陷阱。为了提炼预测的内在质量,我们将预测按近似相等的率分层到不同桶中,并分别评估每个桶的指标。通过将每个桶的达到值与基准进行比较,我们获得了计数预测的缩放感知评级。我们的程序避免了朴素缩放陷阱,提供了对预测质量的直观即时判断,并允许比较不同产品或甚至行业的预测。