Generative AI including large language models (LLMs) have recently gained significant interest in the geo-science community through its versatile task-solving capabilities including coding, spatial computations, generation of sample data, time-series forecasting, toponym recognition, or image classification. So far, the assessment of LLMs for spatial tasks has primarily focused on ChatGPT, arguably the most prominent AI chatbot, whereas other chatbots received less attention. To narrow this research gap, this study evaluates the correctness of responses for a set of 54 spatial tasks assigned to four prominent chatbots, i.e., ChatGPT-4, Bard, Claude-2, and Copilot. Overall, the chatbots performed well on spatial literacy, GIS theory, and interpretation of programming code and given functions, but revealed weaknesses in mapping, code generation, and code translation. ChatGPT-4 outperformed other chatbots across most task categories.
翻译:生成式人工智能,包括大型语言模型(LLMs),最近因其多功能任务解决能力(包括编码、空间计算、样本数据生成、时间序列预测、地名识别或图像分类)而在地球科学界引起了极大关注。到目前为止,对LLMs在空间任务上的评估主要聚焦于ChatGPT(可以说是最著名的人工智能聊天机器人),而其他聊天机器人受到的关注较少。为弥补这一研究空白,本研究评估了四款主流聊天机器人(即ChatGPT-4、Bard、Claude-2和Copilot)对54项空间任务回答的正确性。总体而言,这些聊天机器人在空间素养、GIS理论、编程代码及给定函数解释方面表现良好,但在制图、代码生成和代码翻译方面显示出不足。ChatGPT-4在大多数任务类别中优于其他聊天机器人。