Sensemaking on a large collection of documents (corpus) is a challenging task often found in fields such as market research, legal studies, intelligence analysis, political science, computational linguistics, etc. Previous works approach this problem either from a topic- or entity-based perspective, but they lack interpretability and trust due to poor model alignment. In this paper, we present HINTs, a visual analytics approach that combines topic- and entity-based techniques seamlessly and integrates Large Language Models (LLMs) as both a general NLP task solver and an intelligent agent. By leveraging the extraction capability of LLMs in the data preparation stage, we model the corpus as a hypergraph that matches the user's mental model when making sense of the corpus. The constructed hypergraph is hierarchically organized with an agglomerative clustering algorithm by combining semantic and connectivity similarity. The system further integrates an LLM-based intelligent chatbot agent in the interface to facilitate sensemaking. To demonstrate the generalizability and effectiveness of the HINTs system, we present two case studies on different domains and a comparative user study. We report our insights on the behavior patterns and challenges when intelligent agents are used to facilitate sensemaking. We find that while intelligent agents can address many challenges in sensemaking, the visual hints that visualizations provide are necessary to address the new problems brought by intelligent agents. We discuss limitations and future work for combining interactive visualization and LLMs more profoundly to better support corpus analysis.
翻译:大规模文档集(语料库)的意义建构是市场研究、法律分析、情报分析、政治学、计算语言学等领域常见且富有挑战性的任务。已有研究或从主题视角、或从实体视角出发解决该问题,但因模型对齐性差而缺乏可解释性和可信度。本文提出HINTs,一种融合主题与实体技术,并将大型语言模型(LLMs)同时作为通用自然语言处理任务求解器和智能体的可视化分析方法。在数据准备阶段,我们利用LLMs的提取能力,将语料库建模为符合用户意义建构心理模型的超图。通过结合语义相似性与连接相似性的凝聚聚类算法,对构建的超图进行层次化组织。系统进一步在界面中集成基于LLM的智能对话体,以辅助意义建构。为证明HINTs系统的泛化性与有效性,我们展示了两个不同领域的案例研究及一项对比用户实验,并报告了关于智能体辅助意义建构时行为模式与挑战的洞察。研究发现,尽管智能体能解决意义建构中的诸多难题,但可视化提供的视觉线索对于应对智能体带来的新问题至关重要。我们讨论了更深度融合交互式可视化与LLMs以更好支持语料库分析的局限性及未来工作方向。