Financial information no longer arrives in a single format. Research reports come as PDFs, financial statements live in spreadsheets, market trends are captured in images, and policy documents reach analysts as scans, each carrying part of the picture the others cannot supply. Accounting information systems built around single-modality extraction pipelines and rule-based tools therefore struggle to assemble the full picture, slowing financial statement analysis, complicating audit evidence corroboration, and limiting investment decision support. This study presents FinVision, a multimodal large language model that unites vision-language models with domain-specific financial reasoning. Instead of processing documents in isolation, FinVision reads text, tables, and images together, converts them into consistent structured data, and verifies cross-modal agreement, in the same spirit as auditors corroborating evidence from independent sources. The model is trained in two stages, pre-trained on large-scale public financial corpora and fine-tuned on institution-specific investment data, so it can apply established valuation methodologies and audit risk assessment frameworks while outperforming zero-shot and single-stage baselines. A natural-language decision pipeline lets users describe what they need and turns those descriptions into executable workflows, supporting portfolio optimization, real-time risk monitoring, and refinement through multi-turn dialogue. Across 200 listed companies, FinVision reduced valuation error by 19 percent relative to the strongest baseline; a user study with 48 accounting and investment professionals reported a 51 percent reduction in task completion time. These results carry implications for audit automation, financial reporting quality, and more inclusive access to expert-level financial analysis.
翻译:暂无翻译