Data analysis is challenging as it requires synthesizing domain knowledge, statistical expertise, and programming skills. Assistants powered by large language models (LLMs), such as ChatGPT, can assist analysts by translating natural language instructions into code. However, AI-assistant responses and analysis code can be misaligned with the analyst's intent or be seemingly correct but lead to incorrect conclusions Therefore, validating AI assistance is crucial and challenging. Here, we explore how analysts across a range of backgrounds and expertise understand and verify the correctness of AI-generated analyses. We develop a design probe that allows analysts to pursue diverse verification workflows using natural language explanations, code, visualizations, inspecting data tables, and performing common data operations. Through a qualitative user study (n=22) using this probe, we uncover common patterns of verification workflows influenced by analysts' programming, analysis, and AI backgrounds. Additionally, we highlight open challenges and opportunities for improving future AI analysis assistant experiences.
翻译:数据分析具有挑战性,因为它需要综合领域知识、统计专业知识和编程技能。基于大语言模型(LLM)的助手(如ChatGPT)可通过将自然语言指令转化为代码来协助分析师。然而,AI助手的响应和分析代码可能偏离分析师的意图,或看似正确但实则导出错误结论。因此,验证AI辅助结果的正确性至关重要且充满挑战。本研究探索了不同背景和专业水平的分析师如何理解与验证AI生成分析结果的正确性。我们开发了一个设计探针,使分析师能够通过自然语言解释、代码、可视化、数据表检查及执行常规数据操作等多种验证流程开展工作。通过基于该探针的定性用户研究(n=22),我们揭示了受分析师编程、分析及AI背景影响的常见验证流程模式,并指出了改进未来AI分析助手体验的开放挑战与机遇。