Spatially correlated data with an excess of zeros, usually referred to as zero-inflated spatial data, arise in many disciplines. Examples include count data, for instance, abundance (or lack thereof) of animal species and disease counts, as well as semi-continuous data like observed precipitation. Spatial two-part models are a flexible class of models for such data. Fitting two-part models can be computationally expensive for large data due to high-dimensional dependent latent variables, costly matrix operations, and slow mixing Markov chains. We describe a flexible, computationally efficient approach for modeling large zero-inflated spatial data using the projection-based intrinsic conditional autoregression (PICAR) framework. We study our approach, which we call PICAR-Z, through extensive simulation studies and two environmental data sets. Our results suggest that PICAR-Z provides accurate predictions while remaining computationally efficient. An important goal of our work is to allow researchers who are not experts in computation to easily build computationally efficient extensions to zero-inflated spatial models; this also allows for a more thorough exploration of modeling choices in two-part models than was previously possible. We show that PICAR-Z is easy to implement and extend in popular probabilistic programming languages such as nimble and stan.
翻译:零膨胀空间数据(即存在大量零值且具有空间相关性的数据)在多个学科中普遍出现。例如计数数据(如物种丰度或缺失数据)以及半连续数据(如观测降水量)。空间两部件模型是针对此类数据的一类灵活建模框架。然而,由于高维潜在变量间的依赖关系、高成本的矩阵运算以及马尔可夫链混合缓慢等问题,对大规模数据拟合两部件模型在计算上成本高昂。本文基于投影固有条件自回归(PICAR)框架,提出一种用于大规模零膨胀空间数据建模的灵活高效计算方法。我们通过大量模拟实验及两个环境数据集对所提出的PICAR-Z方法进行验证。结果表明,PICAR-Z在保持计算高效性的同时,能提供精准的预测。本工作的重要目标是让非计算领域的研究人员能够便捷地构建零膨胀空间模型的高效扩展,从而更深入地探索两部件模型中的建模选择。我们证明,PICAR-Z可在nimble和stan等主流概率编程语言中简便实现与扩展。