This paper details a methodology proposed for the EVA 2021 conference data challenge. The aim of this challenge was to predict the number and size of wildfires over the contiguous US between 1993 and 2015, with more importance placed on extreme events. In the data set provided, over 14\% of both wildfire count and burnt area observations are missing; the objective of the data challenge was to estimate a range of marginal probabilities from the distribution functions of these missing observations. To enable this prediction, we make the assumption that the marginal distribution of a missing observation can be informed using non-missing data from neighbouring locations. In our method, we select spatial neighbourhoods for each missing observation and fit marginal models to non-missing observations in these regions. For the wildfire counts, we assume the compiled data sets follow a zero-inflated negative binomial distribution, while for burnt area values, we model the bulk and tail of each compiled data set using non-parametric and parametric techniques, respectively. Cross validation is used to select tuning parameters, and the resulting predictions are shown to significantly outperform the benchmark method proposed in the challenge outline. We conclude with a discussion of our modelling framework, and evaluate ways in which it could be extended.
翻译:本文详细阐述了为EVA 2021会议数据挑战赛提出的方法论。本次挑战赛的目标是预测1993年至2015年间美国本土野火事件的次数与规模,其中极端事件被赋予更高权重。在提供的数据集中,超过14%的野火次数和燃烧面积观测值存在缺失;数据挑战赛的目标是通过这些缺失观测值的分布函数估计一系列边际概率。为实现预测,我们假设缺失观测值的边际分布可借助邻近位置的非缺失数据推断。在方法中,我们为每个缺失观测值选取空间邻域,并对这些区域内的非缺失观测值拟合边际模型。针对野火次数,我们假设汇总数据集遵循零膨胀负二项分布;而针对燃烧面积值,则分别采用非参数技术建模数据集的主体部分,参数技术建模其尾部。通过交叉验证选择调优参数,结果表明所得预测显著优于挑战赛大纲中提出的基准方法。最终,我们讨论了建模框架的局限性,并评估了可能的扩展方向。