Current Large Language Models (LLMs), especially Large Reasoning Models, can generate Chain-of-Thought (CoT) reasoning traces to illustrate how they produce final outputs, thereby facilitating trust calibration for users. However, these CoT reasoning traces are usually lengthy and tedious, and can contain various issues, such as logical and factual errors, which make it difficult for users to interpret the reasoning traces efficiently and accurately. To address these challenges, we develop an error detection pipeline that combines external fact-checking with symbolic formal logical validation to identify errors at the step level. Building on this pipeline, we propose ReasonDiag, an interactive visualization system for diagnosing CoT reasoning traces. ReasonDiag provides 1) an integrated arc diagram to show reasoning-step distributions and error-propagation patterns, and 2) a hierarchical node-link diagram to visualize high-level reasoning flows and premise dependencies. We evaluate ReasonDiag through a technical evaluation for the error detection pipeline, two case studies, and user interviews with 16 participants. The results indicate that ReasonDiag helps users effectively understand CoT reasoning traces, identify erroneous steps, and determine their root causes.
翻译:当前的大语言模型(LLMs),特别是大型推理模型,能够生成思维链(Chain-of-Thought, CoT)推理轨迹,以展示其如何产生最终输出,从而有助于用户建立信任校准。然而,这些CoT推理轨迹通常冗长繁琐,且可能包含逻辑和事实错误等多种问题,导致用户难以高效、准确地解读推理轨迹。为解决这些挑战,我们开发了一个错误检测流水线,该流水线结合了外部事实核查与符号形式逻辑验证,以在步骤级别识别错误。基于此流水线,我们提出了ReasonDiag,一个用于诊断CoT推理轨迹的交互式可视化系统。ReasonDiag提供了:1)集成的弧形图,用于展示推理步骤分布和错误传播模式;2)层级节点-链接图,用于可视化高层推理流程和前提依赖关系。我们通过技术评估(针对错误检测流水线)、两项案例研究以及16名参与者的用户访谈对ReasonDiag进行了评估。结果表明,ReasonDiag能帮助用户有效理解CoT推理轨迹,识别错误步骤并确定其根本原因。