Accurate estimation of long-term care transition probabilities is central to disability insurance pricing, reserving, and solvency assessment. Classical actuarial multi-state models commonly rely on Markov, semi-Markov, or proportional-hazard specifications, which provide a direct connection to cohort projection but may be restrictive for irregular longitudinal health data with nonlinear aging patterns and heterogeneous covariate histories. This paper develops a well-calibrated estimator of multi-state transition probabilities for irregular longitudinal health data. The model learns from individual health history, incorporates the time elapsed between observations, and conditions transition probabilities on demographic and socioeconomic attributes. It produces a valid probability distribution over the next observed health state, with four possible states: healthy, mild disability, severe disability, and death. Individual probabilities are aggregated by age group and origin state to form transition matrices compatible with actuarial cohort projection. Using longitudinal data from the Health and Retirement Study, we compare the proposed estimator with logistic regression, gradient-boosted trees, a recurrent neural network, and a last-state persistence benchmark. The evaluation considers probabilistic accuracy, endpoint discrimination and calibration for severe disability and death, risk concentration, and transition matrix error after aggregation. The proposed estimator improves severe disability discrimination relative to logistic regression and gradient-boosted tree benchmarks, maintains strong calibration, and yields the lowest transition matrix error among the evaluated models in the held-out test analysis. Results show that a structured machine learning estimator can support long-term care transition modeling when judged by calibration and projection fidelity, beyond discrimination.
翻译:长期护理转移概率的精确估计是残疾保险定价、准备金计提与偿付能力评估的核心环节。经典精算多状态模型通常采用马尔可夫、半马尔可夫或比例风险设定,这类模型虽能直接生成队列预测,但对具有非线性老龄化模式及异质性协变量历史的非规则纵向健康数据可能存在约束性限制。本文针对非规则纵向健康数据,开发了一种经过充分校准的多状态转移概率估计方法。该模型通过学习个体健康史信息,整合观测时间间隔,并依据人口统计学特征和社会经济属性对转移概率进行条件建模。模型可生成下一观测健康状态的有效概率分布,涵盖四种可能状态:健康、轻度失能、重度失能与死亡。通过按年龄组和初始状态聚合个体概率,形成与精算队列预测兼容的转移矩阵。基于健康与退休研究的纵向数据,我们将该估计方法与逻辑回归、梯度提升树、循环神经网络及最后状态持久化基准模型进行对比。评估指标涵盖概率准确性、重度失能与死亡状态的端点判别能力与校准度、风险集中度,以及聚合后的转移矩阵误差。相较于逻辑回归和梯度提升树基准模型,本文方法显著提升了重度失能的判别能力,同时保持强校准性,并在留出测试分析中展现了最低的转移矩阵误差。研究结果表明,当以校准度和预测保真度作为评判标准时,结构化机器学习估计方法可有效支持长期护理转移建模,其性能优势不仅局限于判别能力。