Count data with complex features arise in many disciplines, including ecology, agriculture, criminology, medicine, and public health. Zero inflation, spatial dependence, and non-equidispersion are common features in count data. There are two classes of models that allow for these features -- the mode-parameterized Conway--Maxwell--Poisson (COMP) distribution and the generalized Poisson model. However both require the use of either constraints on the parameter space or a parameterization that leads to challenges in interpretability. We propose a spatial mean-parameterized COMP model that retains the flexibility of these models while resolving the above issues. We use a Bayesian spatial filtering approach in order to efficiently handle high-dimensional spatial data and we use reversible-jump MCMC to automatically choose the basis vectors for spatial filtering. The COMP distribution poses two additional computational challenges -- an intractable normalizing function in the likelihood and no closed-form expression for the mean. We propose a fast computational approach that addresses these challenges by, respectively, introducing an efficient auxiliary variable algorithm and pre-computing key approximations for fast likelihood evaluation. We illustrate the application of our methodology to simulated and real datasets, including Texas HPV-cancer data and US vaccine refusal data.
翻译:在生态学、农学、犯罪学、医学和公共卫生等多个学科领域中,常出现具有复杂特征的计数数据。零膨胀、空间依赖性和非等离散性是计数数据的常见特征。目前有两类模型能处理这些特征:参数化众数的Conway–Maxwell–Poisson分布和广义泊松模型。然而,这两类模型要么需要对参数空间施加约束,要么采用导致可解释性问题的参数化方案。我们提出了一种空间均值参数化的COMP模型,在保留上述模型灵活性的同时解决了这些问题。为了高效处理高维空间数据,我们采用贝叶斯空间滤波方法,并利用可逆跳跃马尔可夫链蒙特卡洛方法自动选择空间滤波的基向量。COMP分布带来两个额外的计算挑战:似然函数中存在难以处理的正则化函数,以及均值缺少闭式表达式。我们提出了一种快速计算方法,通过引入高效的辅助变量算法处理前者,并预先计算关键近似值以加速似然估计来处理后者。我们通过模拟数据集和真实数据集(包括德克萨斯州HPV癌症数据与美国疫苗拒接数据)展示了该方法的应用效果。