Data visualization serves as a critical means for presenting data and mining its valuable insights. The task of chart summarization, through natural language processing techniques, facilitates in-depth data analysis of charts. However, there still are notable deficiencies in terms of visual-language matching and reasoning ability for existing approaches. To address these limitations, this study constructs a large-scale dataset of comprehensive chart-caption pairs and fine-tuning instructions on each chart. Thanks to the broad coverage of various topics and visual styles within this dataset, better matching degree can be achieved from the view of training data. Moreover, we propose an innovative chart summarization method, ChartThinker, which synthesizes deep analysis based on chains of thought and strategies of context retrieval, aiming to improve the logical coherence and accuracy of the generated summaries. Built upon the curated datasets, our trained model consistently exhibits superior performance in chart summarization tasks, surpassing 8 state-of-the-art models over 7 evaluation metrics. Our dataset and codes are publicly accessible.
翻译:数据可视化是呈现数据并挖掘其宝贵见解的关键手段。图表摘要任务通过自然语言处理技术促进对图表的深度数据分析。然而,现有方法在视觉-语言匹配和推理能力方面仍存在显著不足。为解决这些局限,本研究构建了一个涵盖全面图表-描述配对及每个图表微调指令的大规模数据集。得益于该数据集对各类主题和视觉风格的广泛覆盖,可从训练数据角度实现更好的匹配度。此外,我们提出了一种创新的图表摘要方法ChartThinker,该方法基于思维链和上下文检索策略合成深度分析,旨在提升生成摘要的逻辑连贯性和准确性。基于所整理的数据集,我们训练的模型在图表摘要任务中始终表现出卓越性能,在7项评估指标上超越了8个最先进的模型。我们的数据集和代码已公开。