High-cardinality categorical features are pervasive in actuarial data (e.g. occupation in commercial property insurance). Standard categorical encoding methods like one-hot encoding are inadequate in these settings. In this work, we present a novel _Generalised Linear Mixed Model Neural Network_ ("GLMMNet") approach to the modelling of high-cardinality categorical features. The GLMMNet integrates a generalised linear mixed model in a deep learning framework, offering the predictive power of neural networks and the transparency of random effects estimates, the latter of which cannot be obtained from the entity embedding models. Further, its flexibility to deal with any distribution in the exponential dispersion (ED) family makes it widely applicable to many actuarial contexts and beyond. We illustrate and compare the GLMMNet against existing approaches in a range of simulation experiments as well as in a real-life insurance case study. Notably, we find that the GLMMNet often outperforms or at least performs comparably with an entity embedded neural network, while providing the additional benefit of transparency, which is particularly valuable in practical applications. Importantly, while our model was motivated by actuarial applications, it can have wider applicability. The GLMMNet would suit any applications that involve high-cardinality categorical variables and where the response cannot be sufficiently modelled by a Gaussian distribution.
翻译:高基数分类特征在精算数据中普遍存在(例如商业财产保险中的职业)。标准的分类编码方法如独热编码在这些场景中并不适用。在本工作中,我们提出了一种新颖的"广义线性混合模型神经网络"("GLMMNet")方法,用于对高基数分类特征进行建模。GLMMNet将广义线性混合模型集成到深度学习框架中,既提供了神经网络的预测能力,又保留了随机效应估计的可解释性——后者无法从实体嵌入模型中获取。此外,其灵活处理指数分散(ED)族中任意分布的能力,使其广泛适用于众多精算应用场景及其他领域。我们通过一系列模拟实验以及一项真实保险案例研究,将GLMMNet与现有方法进行了对比。值得注意的是,我们发现GLMMNet通常优于或至少能与实体嵌入神经网络表现相当,同时额外提供了可解释性的优势,这在实际应用中尤为珍贵。重要的是,尽管我们的模型源于精算应用需求,但其具有更广泛的适用性。GLMMNet适用于任何涉及高基数分类变量且响应变量无法通过高斯分布充分建模的应用场景。