Count data with excess zeros arise frequently in health economics and epidemiology. The standard Poisson Hurdle Model (PHM) parametrises the underlying Poisson rate directly, so its count-component coefficients are log-rate ratios rather than log-ratios of the marginal mean. Consequently, the incidence density ratio (IDR) from the PHM is neither exact nor constant across covariate profiles, complicating applied reporting. We propose the Marginalised Poisson Hurdle Model (MPHM), which reparametrises the count component so that the coefficient vector beta directly governs the marginal mean E[Y]. A nonlinear connector equation links the structural Poisson rate to this parametrised mean. We prove existence and uniqueness of the connector solution, develop a vectorised Brent's-method solver, derive the score equations and block-diagonal Fisher information, establish asymptotic normality, and prove that exp(beta) is exactly constant across all covariate values. A simulation study with n in {100, 250, 500, 1000}, zero proportion pi in {0.2, 0.4, 0.6, 0.8}, and R = 200 replications confirms consistency, near-zero bias, and 95% Wald coverage of 0.905-0.975 across all 16 scenarios. Applied to the NMES1988 physician visit data (n = 4,406), the MPHM yields IDR = 1.163 (95% CI: 1.150-1.177) per additional chronic condition - an exact, population-wide effect not derivable from the PHM. The MPHM resolves the non-constant IDR problem by directly parametrising E[Y]. The resulting IDR holds for every individual and the whole population without further marginalisation, substantially simplifying the reporting of covariate effects in health utilisation research.


翻译:含过量零值的计数数据在健康经济学与流行病学中频繁出现。标准泊松障碍模型(PHM)直接参数化潜在泊松率,其计数分量的系数为对数率比而非边际均值的对数比,导致PHM的发病密度比(IDR)既不精确也不随协变量分布恒定,给应用报告带来困难。我们提出边际化泊松障碍模型(MPHM),通过重新参数化计数分量,使系数向量β直接控制边际均值E[Y],并借助非线性连接方程将结构泊松率与该参数化均值关联。我们证明了连接解的存在唯一性,开发了向量化Brent法求解器,推导了得分方程与块对角Fisher信息矩阵,建立了渐近正态性,并证明exp(β)对所有协变量取值严格恒定。基于n∈{100,250,500,1000}、零比例π∈{0.2,0.4,0.6,0.8}及R=200次重复的模拟研究证实:在全部16种场景下,模型具有一致性、近零偏差,且95% Wald覆盖率介于0.905-0.975。应用于NMES1988医生就诊数据(n=4,406)时,MPHM显示每增加一种慢性病,IDR=1.163(95%CI: 1.150-1.177)——该精确且具群体普适性的效应无法通过PHM推导。MPHM通过直接参数化E[Y]解决了非恒定IDR问题,所得IDR对每个个体及全体人群均成立且无需额外边际化,显著简化了健康利用研究中协变量效应的报告。

0
下载
关闭预览

相关内容

大型语言模型的规模效应局限
专知会员服务
14+阅读 · 2025年11月18日
零样本量化:综述
专知会员服务
13+阅读 · 2025年5月15日
不平衡数据学习的全面综述
专知会员服务
44+阅读 · 2025年2月15日
【NeurIPS2024】用于缺失值数据集的可解释广义加性模型
专知会员服务
18+阅读 · 2024年12月7日
【NeurIPS2024】通过方差减少实现零样本模型的稳健微调
专知会员服务
19+阅读 · 2024年11月12日
用于疾病诊断的大型语言模型:范围综述
专知会员服务
26+阅读 · 2024年9月8日
【CMU博士论文】分布偏移下的不确定性量化,226页pdf
专知会员服务
31+阅读 · 2023年9月30日
《过参数化机器学习理论》综述论文
专知会员服务
46+阅读 · 2021年9月19日
【斯坦福经典书】统计学稀疏性:Lasso与泛化性,362页pdf
专知会员服务
37+阅读 · 2020年11月15日
【边缘计算】边缘计算面临的问题
产业智能官
17+阅读 · 2019年5月31日
换个角度看GAN:另一种损失函数
机器之心
16+阅读 · 2019年1月1日
数据分析师应该知道的16种回归方法:泊松回归
数萃大数据
35+阅读 · 2018年9月13日
超全总结:神经网络加速之量化模型 | 附带代码
概率论之概念解析:边缘化(Marginalisation)
边缘计算:万物互联时代新型计算模型
计算机研究与发展
15+阅读 · 2017年5月19日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
对抗环境下超视距目标打击的情报支援
专知会员服务
3+阅读 · 今天14:49
《无人机对海面作战影响评估》
专知会员服务
11+阅读 · 7月21日
印度精确打击与指挥架构的断层
专知会员服务
6+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
6+阅读 · 7月20日
相关VIP内容
大型语言模型的规模效应局限
专知会员服务
14+阅读 · 2025年11月18日
零样本量化:综述
专知会员服务
13+阅读 · 2025年5月15日
不平衡数据学习的全面综述
专知会员服务
44+阅读 · 2025年2月15日
【NeurIPS2024】用于缺失值数据集的可解释广义加性模型
专知会员服务
18+阅读 · 2024年12月7日
【NeurIPS2024】通过方差减少实现零样本模型的稳健微调
专知会员服务
19+阅读 · 2024年11月12日
用于疾病诊断的大型语言模型:范围综述
专知会员服务
26+阅读 · 2024年9月8日
【CMU博士论文】分布偏移下的不确定性量化,226页pdf
专知会员服务
31+阅读 · 2023年9月30日
《过参数化机器学习理论》综述论文
专知会员服务
46+阅读 · 2021年9月19日
【斯坦福经典书】统计学稀疏性:Lasso与泛化性,362页pdf
专知会员服务
37+阅读 · 2020年11月15日
相关基金
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员