It is well-established that the provenance of a scientific result is important, sometimes more important than the actual result. For computational analyses that involve visualization, this provenance information may contain the steps involved in generating visualizations from raw data. Specifically, data provenance tracks the lineage of data and process provenance tracks the steps executed. In this paper, we argue that the utility of computational provenance may not be as clear-cut as we might like. One common use case for provenance is that the information can be used to reproduce the original result. However, in visualization, the goal is often to communicate results to a user or viewer, and thus the insights obtained are ultimately most important. Viewers can miss important changes or react to unimportant ones. Here, interaction provenance, which tracks a user's actions with a visualization, or insight provenance, which tracks the decision-making process, can help capture what happened but don't remove the issues. In this paper, we present scenarios where provenance impacts reproducibility in different ways. We also explore how provenance and visualizations can be better related.
翻译:科学结果的来源信息至关重要,有时甚至比实际结果本身更重要,这一点已得到广泛认可。对于涉及可视化的计算分析而言,此类来源信息可能包含从原始数据生成可视化结果所涉及的步骤。具体而言,数据来源追踪数据的谱系,而过程来源追踪已执行的步骤。本文提出,计算来源信息的效用可能并非如我们期望的那样明确。来源信息的一个常见用途是用于再现原始结果。然而,在可视化领域,目标往往是向用户或观看者传达结果,因此最终获取的洞察最为关键。观看者可能会遗漏重要变化,或对无关紧要的变化做出反应。在此情境下,交互来源(追踪用户与可视化交互的行为)或洞察来源(追踪决策过程)有助于捕捉已发生的事件,但无法消除上述问题。本文展示了来源信息以不同方式影响可重复性的场景,并探讨了如何更好地关联来源信息与可视化。