In this paper, we study mid-cap companies, i.e. publicly traded companies with less than US $10 billion in market capitalisation. Using a large dataset of US mid-cap companies observed over 30 years, we look to predict the default probability term structure over the medium term and understand which data sources (i.e. fundamental, market or pricing data) contribute most to the default risk. Whereas existing methods typically require that data from different time periods are first aggregated and turned into cross-sectional features, we frame the problem as a multi-label time-series classification problem. We adapt transformer models, a state-of-the-art deep learning model emanating from the natural language processing domain, to the credit risk modelling setting. We also interpret the predictions of these models using attention heat maps. To optimise the model further, we present a custom loss function for multi-label classification and a novel multi-channel architecture with differential training that gives the model the ability to use all input data efficiently. Our results show the proposed deep learning architecture's superior performance, resulting in a 13% improvement in AUC (Area Under the receiver operating characteristic Curve) over traditional models. We also demonstrate how to produce an importance ranking for the different data sources and the temporal relationships using a Shapley approach specific to these models.
翻译:本文研究中市值公司,即市值低于100亿美元的上市公司。通过使用涵盖30年美国中市值公司的大规模数据集,我们旨在预测其中期违约概率期限结构,并探究何种数据源(基本面数据、市场数据或定价数据)对违约风险贡献最大。现有方法通常需要先对不同时期数据进行聚合处理并转化为横截面特征,而我们将此问题重构为多标签时间序列分类任务。我们借鉴自然语言处理领域的先进深度学习模型Transformer,将其改进后应用于信用风险建模场景,并利用注意力热力图解读模型预测结果。为进一步优化模型,我们提出专门用于多标签分类的自定义损失函数,以及采用差异化训练的新型多通道架构,使模型能够高效利用所有输入数据。实验结果表明,所提出的深度学习架构性能显著优于传统模型,在AUC(受试者工作特征曲线下面积)指标上实现了13%的提升。我们还展示了如何运用专为这些模型设计的Shapley方法,对不同数据源的重要性及时间依赖关系进行排序。