We introduce FinWorkBench (a.k.a. Finch), a benchmark for evaluating agents on real-world, enterprise-grade finance and accounting workflows that interleave data entry, structuring, formatting, web search, cross-file retrieval, calculation, modeling, validation, translation, visualization, and reporting. Finch is built from authentic enterprise workspaces from Enron (15,000 files and 500,000 emails) and other financial institutions spanning 2000 to 2025, preserving the in-the-wild messiness of multimodal artifacts such as tables and charts across diverse domains including budgeting, trading, and asset management. We propose a workflow construction process that combines LLM-assisted mining of workflows from authentic enterprise environments with expert annotation. Specifically, we use LLM-assisted, expert-verified derivation of workflows from real-world email threads and spreadsheet version histories, followed by meticulous workflow annotation requiring more than 700 hours of expert effort. This process yields 172 composite workflows with 384 tasks, involving 1,710 spreadsheets with 27 million cells, along with PDFs and other artifacts, capturing the intrinsically messy, long-horizon, knowledge-intensive, and collaborative nature of enterprise work. We conduct both human and automated evaluations of frontier AI systems, including GPT 5.1, Claude Sonnet/Opus 4.5, Gemini 3 Pro, Grok 4, and Qwen 3 Max. GPT 5.1 Pro spends an average of 16.8 minutes per workflow yet passes only 38.4% of workflows. Comprehensive case studies further highlight the challenges that real-world enterprise workflows pose for AI agents.
翻译:摘要:我们提出FinWorkBench(简称Finch),这是一个用于评估智能体在真实企业级金融与会计工作流中表现的基准测试,涵盖数据录入、结构化处理、格式编排、网络搜索、跨文件检索、计算建模、验证、翻译、可视化及报告生成等环节。Finch基于安然公司(包含15,000份文件和500,000封电子邮件)及其他金融机构在2000年至2025年间真实的业务工作空间构建,保留了跨预算编制、交易、资产管理等多领域的多模态制品(如表格、图表)的自然杂散特性。我们提出一种结合大语言模型辅助的真实企业环境工作流挖掘与专家标注的工作流构建流程。具体而言,我们通过大语言模型辅助、专家验证的方式,从真实邮件线程和电子表格版本历史中推导工作流,随后进行耗时超过700小时的精细标注。该流程最终生成172个复合工作流(含384个任务),涉及1,710份电子表格(共2,700万单元格)及PDF等其他文件,真实捕捉了企业工作固有的杂乱性、长周期、知识密集型和协作性特征。我们对GPT 5.1、Claude Sonnet/Opus 4.5、Gemini 3 Pro、Grok 4及Qwen 3 Max等前沿AI系统进行了人工与自动化评估。GPT 5.1 Pro在每个工作流上平均耗时16.8分钟,但仅通过38.4%的工作流。综合案例研究进一步揭示了真实企业工作流对AI智能体构成的挑战。