A composite likelihood is an inference function derived by multiplying a set of likelihood components. This approach provides a flexible framework for drawing inference when the likelihood function of a statistical model is computationally intractable. While composite likelihood has computational advantages, it can still be demanding when dealing with numerous likelihood components and a large sample size. This paper tackles this challenge by employing an approximation of the conventional composite likelihood estimator, which is derived from an optimization procedure relying on stochastic gradients. This novel estimator is shown to be asymptotically normally distributed around the true parameter. In particular, based on the relative divergent rate of the sample size and the number of iterations of the optimization, the variance of the limiting distribution is shown to compound for two sources of uncertainty: the sampling variability of the data and the optimization noise, with the latter depending on the sampling distribution used to construct the stochastic gradients. The advantages of the proposed framework are illustrated through simulation studies on two working examples: an Ising model for binary data and a gamma frailty model for count data. Finally, a real-data application is presented, showing its effectiveness in a large-scale mental health survey.
翻译:复合似然是一种通过将一组似然分量相乘而得到的推断函数。当统计模型的对数似然函数在计算上难以处理时,该方法为推断提供了灵活的框架。尽管复合似然具有计算优势,但在处理大量似然分量和大样本量时仍可能面临挑战。本文通过采用依赖于随机梯度的优化过程推导出的传统复合似然估计量的近似形式来应对这一挑战。该新型估计量被证明能够渐近地收敛到真实参数的正态分布。特别地,基于样本量与优化迭代次数的相对发散速率,极限分布的方差被证明包含两类不确定性来源:数据的抽样变异性与优化噪声,其中后者依赖于用于构造随机梯度的采样分布。通过两个实例的模拟研究——二元数据的伊辛模型和计数数据的伽马脆弱性模型——展示了所提框架的优势。最后,通过一项大规模心理健康调查的实际数据应用,验证了该方法在真实场景中的有效性。