Financial risk prediction plays a crucial role in the financial sector. Machine learning methods have been widely applied for automatically detecting potential risks and thus saving the cost of labor. However, the development in this field is lagging behind in recent years by the following two facts: 1) the algorithms used are somewhat outdated, especially in the context of the fast advance of generative AI and large language models (LLMs); 2) the lack of a unified and open-sourced financial benchmark has impeded the related research for years. To tackle these issues, we propose FinPT and FinBench: the former is a novel approach for financial risk prediction that conduct Profile Tuning on large pretrained foundation models, and the latter is a set of high-quality datasets on financial risks such as default, fraud, and churn. In FinPT, we fill the financial tabular data into the pre-defined instruction template, obtain natural-language customer profiles by prompting LLMs, and fine-tune large foundation models with the profile text to make predictions. We demonstrate the effectiveness of the proposed FinPT by experimenting with a range of representative strong baselines on FinBench. The analytical studies further deepen the understanding of LLMs for financial risk prediction.
翻译:金融风险预测在金融领域至关重要。机器学习方法已被广泛应用于自动检测潜在风险,从而节约人力成本。然而,近年来该领域的发展受到以下两个事实的制约:1)所使用的算法略显过时,尤其在生成式人工智能与大型语言模型(LLMs)快速发展的背景下;2)缺乏统一且开源的金融基准数据集长期阻碍了相关研究。为解决这些问题,我们提出FinPT与FinBench:前者是一种新颖的金融风险预测方法,通过对大型预训练基础模型进行画像调优(Profile Tuning);后者是一组高质量的金融风险数据集,涵盖违约、欺诈及客户流失等场景。在FinPT中,我们将金融表格数据填充至预定义的指令模板,通过提示LLMs生成自然语言形式的客户画像,并利用画像文本微调大型基础模型以实现预测。通过在FinBench上对比一系列代表性的强基线方法,我们验证了所提FinPT的有效性。分析性研究进一步加深了对LLMs在金融风险预测中作用的理解。