In the financial industry, credit scoring is a fundamental element, shaping access to credit and determining the terms of loans for individuals and businesses alike. Traditional credit scoring methods, however, often grapple with challenges such as narrow knowledge scope and isolated evaluation of credit tasks. Our work posits that Large Language Models (LLMs) have great potential for credit scoring tasks, with strong generalization ability across multiple tasks. To systematically explore LLMs for credit scoring, we propose the first open-source comprehensive framework. We curate a novel benchmark covering 9 datasets with 14K samples, tailored for credit assessment and a critical examination of potential biases within LLMs, and the novel instruction tuning data with over 45k samples. We then propose the first Credit and Risk Assessment Large Language Model (CALM) by instruction tuning, tailored to the nuanced demands of various financial risk assessment tasks. We evaluate CALM, and existing state-of-art (SOTA) open source and close source LLMs on the build benchmark. Our empirical results illuminate the capability of LLMs to not only match but surpass conventional models, pointing towards a future where credit scoring can be more inclusive, comprehensive, and unbiased. We contribute to the industry's transformation by sharing our pioneering instruction-tuning datasets, credit and risk assessment LLM, and benchmarks with the research community and the financial industry.
翻译:在金融行业中,信用评分是核心要素,决定着个人和企业的信贷获取资格与贷款条款。然而,传统信用评分方法常面临知识范围狭窄、信用任务评估孤立等挑战。本研究认为大语言模型(LLMs)在信用评分任务中具有巨大潜力,能够跨多个任务实现强大的泛化能力。为系统探索LLMs在信用评分中的应用,我们提出了首个开源综合框架。我们构建了一个包含9个数据集、1.4万样本的全新基准,专用于信用评估与大语言模型潜在偏见的批判性检验,并开发了含4.5万样本的创新指令微调数据。在此基础上,我们通过指令微调提出了首个信用与风险评估大语言模型(CALM),专门适配各类金融风险评估任务的精细化需求。我们在构建的基准上评估了CALM及现有最先进的开源与闭源大语言模型。实证结果表明,大语言模型不仅能够媲美甚至超越传统模型,预示着未来信用评分可能更具包容性、全面性和无偏性。通过向学术界和金融业界共享我们开创性的指令微调数据集、信用风险评估大语言模型及基准,我们助力行业的变革转型。