Analysts frequently need to create visualizations in the data analysis process to obtain and communicate insights. To reduce the burden of creating visualizations, previous research has developed various approaches for analysts to create visualizations from natural language queries. Recent studies have demonstrated the capabilities of large language models in natural language understanding and code generation tasks. The capabilities imply the potential of using large language models to generate visualization specifications from natural language queries. In this paper, we evaluate the capability of a large language model to generate visualization specifications on the task of natural language to visualization (NL2VIS). More specifically, we have opted for GPT-3.5 and Vega-Lite to represent large language models and visualization specifications, respectively. The evaluation is conducted on the nvBench dataset. In the evaluation, we utilize both zero-shot and few-shot prompt strategies. The results demonstrate that GPT-3.5 surpasses previous NL2VIS approaches. Additionally, the performance of few-shot prompts is higher than that of zero-shot prompts. We discuss the limitations of GPT-3.5 on NL2VIS, such as misunderstanding the data attributes and grammar errors in generated specifications. We also summarized several directions, such as correcting the ground truth and reducing the ambiguities in natural language queries, to improve the NL2VIS benchmark.
翻译:数据分析师在数据分析过程中常需创建可视化以获取和传达洞见。为减轻创建可视化的负担,已有研究开发了多种方法,使分析师能通过自然语言查询生成可视化。近年研究表明,大语言模型在自然语言理解和代码生成任务中展现出卓越能力,这暗示了利用大语言模型从自然语言查询生成可视化规范的潜力。本文评估了大语言模型在自然语言到可视化(NL2VIS)任务中生成可视化规范的能力。具体而言,我们选用GPT-3.5和Vega-Lite分别作为大语言模型和可视化规范的代表,并在nvBench数据集上开展评估。评估中采用零样本和少样本两种提示策略。结果表明GPT-3.5优于以往NL2VIS方法,且少样本提示的性能高于零样本提示。我们讨论了GPT-3.5在NL2VIS上的局限性,例如数据属性误读及生成规范中的语法错误,并总结了改进NL2VIS基准的若干方向,如修正基准真相和降低自然语言查询歧义性。