Multimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal understanding and generation capabilities. However, their understanding of synthetic charts is limited, while existing benchmarks are simplistic and the charts deviate significantly from real-world examples, making it challenging to accurately assess MLLMs' chart comprehension abilities. Hence, a challenging benchmark is essential for investigating progress and uncovering the limitations of current MLLMs on chart data. In this work, we propose to examine chart comprehension through more complex visual logic and introduce ChartBench, a comprehensive chart benchmark to accurately measure MLLMs' fundamental chart comprehension and data reliability. Specifically, ChartBench consists of \textbf{41} categories, \textbf{2K} charts, and \textbf{16K} QA annotations. While significantly expanding chart types, ChartBench avoids direct labelling of data points, which requires MLLMs to infer values akin to humans by leveraging elements like color, legends, and coordinate systems. We also introduce an improved metric, \textit{Acc+}, which accurately reflects MLLMs' chart comprehension abilities while avoiding labor-intensive manual evaluations or costly GPT-based evaluations. We conduct evaluations on \textbf{12} mainstream open-source models and \textbf{2} outstanding proprietary models. Through extensive experiments, we reveal the limitations of MLLMs on charts and provide insights to inspire the community to pay closer attention to MLLMs' chart comprehension abilities. The benchmark and code will be publicly available for research.
翻译:多模态大语言模型(MLLMs)展现出卓越的多模态理解与生成能力,但其对合成图表的理解仍存在局限。现有基准测试过于简单且图表与真实场景偏差较大,难以准确评估MLLMs的图表理解能力。因此,亟需建立具有挑战性的基准来探究当前MLLMs在图表数据上的进展与局限性。本文通过更复杂的视觉逻辑来检验图表理解能力,提出ChartBench——一个全面的图表基准测试,用于精确衡量MLLMs的基础图表理解能力与数据可靠性。具体而言,ChartBench包含\textbf{41}个类别、\textbf{2K}张图表和\textbf{16K}条问答标注。在显著扩展图表类型的同时,ChartBench避免直接标注数据点,要求MLLMs通过颜色、图例和坐标系等元素像人类一样推断数值。我们还引入改进指标\textit{Acc+},既能准确反映MLLMs的图表理解能力,又无需耗时的人工评估或昂贵的GPT评估。我们针对\textbf{12}个主流开源模型和\textbf{2}个卓越商业模型进行了评估。通过大量实验,我们揭示了MLLMs在图表方面的局限性,并为学界关注MLLMs的图表理解能力提供启示。该基准测试与代码将公开发布供研究使用。