LLMs have transformed NLP and shown promise in various fields, yet their potential in finance is underexplored due to a lack of thorough evaluations and the complexity of financial tasks. This along with the rapid development of LLMs, highlights the urgent need for a systematic financial evaluation benchmark for LLMs. In this paper, we introduce FinBen, the first comprehensive open-sourced evaluation benchmark, specifically designed to thoroughly assess the capabilities of LLMs in the financial domain. FinBen encompasses 35 datasets across 23 financial tasks, organized into three spectrums of difficulty inspired by the Cattell-Horn-Carroll theory, to evaluate LLMs' cognitive abilities in inductive reasoning, associative memory, quantitative reasoning, crystallized intelligence, and more. Our evaluation of 15 representative LLMs, including GPT-4, ChatGPT, and the latest Gemini, reveals insights into their strengths and limitations within the financial domain. The findings indicate that GPT-4 leads in quantification, extraction, numerical reasoning, and stock trading, while Gemini shines in generation and forecasting; however, both struggle with complex extraction and forecasting, showing a clear need for targeted enhancements. Instruction tuning boosts simple task performance but falls short in improving complex reasoning and forecasting abilities. FinBen seeks to continuously evaluate LLMs in finance, fostering AI development with regular updates of tasks and models.
翻译:大型语言模型(LLMs)已推动自然语言处理领域变革并在多个领域展现出应用潜力,然而由于缺乏系统性评估以及金融任务的复杂性,其在金融领域的潜力尚未得到充分挖掘。随着LLMs的快速迭代发展,迫切需要构建面向金融领域的系统化评估基准。本文提出FinBen——首个专为全面评估LLMs金融领域能力而设计的开源综合评估基准。该基准依据卡特-霍恩-卡罗尔认知能力理论,整合23项金融任务中的35个数据集,按难度梯度分为三个层次,系统评估LLMs在归纳推理、联想记忆、数量推理、晶体智力等方面的认知能力。通过对GPT-4、ChatGPT及最新Gemini等15个代表性LLMs的评估,揭示了其在金融领域的优势与局限。研究显示:GPT-4在量化分析、信息提取、数值推理和股票交易领域表现领先,Gemini则在文本生成与预测任务中表现突出;但两者均难以应对复杂信息提取与预测任务,亟需针对性优化。指令微调虽能提升简单任务性能,却未能显著改进复杂推理与预测能力。FinBench致力于通过持续更新的任务与模型体系,为金融领域LLMs提供动态评估方案,推动人工智能发展。