This paper examines how different types of large language model (LLM) agents perform on scientific visualization (SciVis) tasks, where users generate visualization workflows from natural-language instructions. We compare three primary interaction paradigms, including domain-specific agents with structured tool use, computer-use agents, and general-purpose coding agents, by evaluating eight representative agents across 15 benchmark tasks and measuring visualization quality, efficiency, robustness, and computational cost. We further analyze interaction modalities, including code scripts and model context protocol (MCP) or API calls for structured tool use, as well as command-line interfaces (CLI) and graphical user interfaces (GUI) for more general interaction, while additionally studying the effect of persistent memory in selected agents. The results reveal clear tradeoffs across paradigms and modalities. General-purpose coding agents achieve the highest task success rates but are computationally expensive, while domain-specific agents are more efficient and stable but less flexible. Computer-use agents perform well on individual steps but struggle with longer multi-step workflows, indicating that long-horizon planning is their primary limitation. Across both CLI- and GUI-based settings, persistent memory improves performance over repeated trials, although its benefits depend on the underlying interaction mode and the quality of feedback. These findings suggest that no single approach is sufficient, and future SciVis systems should combine structured tool use, interactive capabilities, and adaptive memory mechanisms to balance performance, robustness, and flexibility.


翻译:本文探究了不同类型的基于大语言模型的智能体在科学可视化任务中的表现,用户通过自然语言指令生成可视化工作流。我们比较了三种主要交互范式:具备结构化工具使用的领域专用智能体、计算机使用智能体和通用编码智能体。通过评估15个基准任务中8个代表性智能体,并衡量可视化质量、效率、鲁棒性和计算成本,我们进一步分析了交互模式,包括用于结构化工具使用的代码脚本、模型上下文协议或API调用,以及用于更通用交互的命令行界面和图形用户界面,同时研究了选定智能体中持久性记忆的影响。结果揭示了不同范式和模式间的明确权衡:通用编码智能体获得最高任务成功率,但计算成本高昂;领域专用智能体更高效稳定,但灵活性不足;计算机使用智能体在单步骤上表现良好,但在较长多步骤工作流中表现欠佳,表明长期规划是其首要限制。在基于CLI和GUI的设置中,持久性记忆均能提升重复试验的表现,但其效益取决于底层交互模式与反馈质量。这些发现表明,单一方法并不足以应对所有情况,未来的科学可视化系统应结合结构化工具使用、交互能力与自适应记忆机制,以平衡性能、鲁棒性和灵活性。

0
下载
关闭预览

相关内容

【EPFL博士论文】大型语言模型时代的协作式智能体
专知会员服务
37+阅读 · 2025年5月16日
基于大语言模型的智能体优化研究综述
专知会员服务
66+阅读 · 2025年3月25日
大语言模型智能体
专知会员服务
101+阅读 · 2024年12月25日
设计和构建强大的大语言模型智能体
专知会员服务
56+阅读 · 2024年10月6日
基于大型语言模型的软件工程智能体综述
专知会员服务
62+阅读 · 2024年9月6日
【深度语义匹配模型】原理篇二:交互篇
AINLP
16+阅读 · 2020年5月18日
NLP实践:对话系统技术原理和应用
AI100
34+阅读 · 2019年3月20日
知识在检索式对话系统的应用
微信AI
32+阅读 · 2018年9月20日
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
7+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
Arxiv
0+阅读 · 6月12日
VIP会员
最新内容
致命七类无人机:无人机时代的演进型合成兵种
《异构无人水面艇集群作战自主制导算法》130页
相关VIP内容
相关基金
国家自然科学基金
3+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
7+阅读 · 2014年12月31日
国家自然科学基金
10+阅读 · 2013年12月31日
国家自然科学基金
18+阅读 · 2009年12月31日
Top
微信扫码咨询专知VIP会员