Credit scorecards are models used for the modelling of the probability of default of clients. The decision to extend credit to an applicant, as well as the price of the credit, is often based on these models. In order to ensure that scorecards remain accurate over time, the hypothesis of population stability is tested periodically; that is, the hypothesis that the distributions of the attributes of clients at the time when the scorecard was developed is still representative of these distributions at review is tested. A number of measures of population stability are used in practice, with several being proposed in the recent literature. This paper provides a critical review of several testing procedures for the mentioned hypothesis. The widely used population stability index is discussed alongside two recently proposed techniques. Additionally, the use of classical goodness-of-fit techniques is considered and the problems associated with large samples are investigated. In addition to the existing testing procedures, we propose two new techniques which can be used to test population stability. The first is based on the calculation of effect sizes which does not suffer the same problems as classical goodness-of-fit techniques when faced with large samples. The second proposed procedure is the so-called overlapping statistic. We argue that this simple measure can be useful due to its intuitive interpretation. In order to demonstrate the use of the various measures, as well as to highlight their strengths and weaknesses, several numerical examples are included.
翻译:信用评分卡是用于建模客户违约概率的模型。是否向申请人提供信贷以及信贷定价通常基于这些模型。为确保评分卡长期保持准确性,需定期检验群体稳定性假设,即检验评分卡开发时的客户属性分布是否仍能代表当前审查时的分布。实践中采用多种群体稳定性度量指标,近期文献中也提出了若干新方法。本文对上述假设的若干检验方法进行了批判性评述。在讨论广泛使用的群体稳定性指数的同时,还分析了两种近期提出的技术。此外,本文考虑了经典拟合优度技术的应用,并探讨了大样本场景下的相关问题。除现有检验方法外,我们提出了两种可用于检验群体稳定性的新技术。第一种基于效应量计算,该方法在面对大样本时不会遭遇经典拟合优度技术的同类问题。第二种提出的方法是所谓的重叠统计量。我们认为,这一简单度量因其直观的解释性而具有实用价值。为展示各度量方法的使用并突出其优劣,文中包含若干数值示例。