Using standard financial ratios as variables in statistical analyses has been related to several serious problems, such as extreme outliers, asymmetry, non-normality, and non-linearity. The compositional-data methodology has been successfully applied to solve these problems and has always yielded substantially different results when compared to standard financial ratios. An under-researched area is the use of financial log-ratios computed with the compositional-data methodology to predict bankruptcy or the related terms of business default, insolvency or failure. Another under-researched area is the use of machine learning methods in combination with compositional log-ratios. The present article adapts the classical Altman bankruptcy prediction model and some of its extensions to the compositional methodology with pairwise log-ratios and three common statistical and machine learning tools: logistic regression models, k-nearest neighbours, and random forests, and compares the results with standard financial ratios. Data from the sector in the Spanish economy with the largest number of bankrupt firms according to the first two digits of the NACE code (46XX "wholesale trade, except of motor vehicles and motorcycles") were obtained from the Iberian Balance sheet Analysis System. The sample size (31,131 firms, of which 97 were bankrupt) was divided into a training and a validation dataset. The training dataset was downsampled to one healthy firm to each bankrupt firm. No outliers were removed. Focusing on predictive performance, the results show that compositional methods are better than standard ratios in terms of sensitivity (recall), with mixed results regarding specificity, compositional random forests and compositional logistic regression behaving the best.
翻译:在统计分析中使用标准财务比率作为变量会引发若干严重问题,例如极端异常值、非对称性、非正态性和非线性。成分数据分析方法已成功应用于解决这些问题,且与标准财务比率相比始终得出显著不同的结果。目前研究不足的领域包括:使用基于成分数据分析方法计算的财务对数比率预测破产或相关术语(如企业违约、资不抵债或经营失败),以及将机器学习方法与成分对数比率结合应用。本文针对经典Altman破产预测模型及其部分扩展模型,采用成分方法论中的配对对数比率,结合三种常见统计与机器学习工具(逻辑回归模型、k近邻算法和随机森林)进行改编,并将结果与标准财务比率进行对比。研究数据来源于西班牙经济中破产企业数量最多的行业(根据NACE代码前两位46XX"批发贸易,机动车和摩托车除外"),取自伊比利亚资产负债表分析系统。样本量为31,131家企业(其中97家破产),被划分为训练集和验证集。训练集通过降采样使健康企业与破产企业比例保持为1:1,未剔除任何异常值。在预测性能方面,结果表明:成分方法在灵敏度(召回率)上优于标准比率,在特异性方面表现参差不齐,其中成分随机森林和成分逻辑回归效果最佳。