Data provenance (the process of determining the origin and derivation of data outputs) has applications across multiple domains including explaining database query results and auditing scientific workflows. Despite decades of research, provenance tracing remains challenging due to its high computational cost and storage requirements. In streaming systems such as Apache Flink, fine-grained provenance graphs can grow super-linearly with data volume, posing significant scalability challenges. We define temporal attribution, a new lightweight form of provenance, appropriate for certain tasks, such as monitoring dependencies between system components over time quantitatively. Temporal attribution enables time-focused analysis that does not require fine-grained, tuple-level dependency meta-data. Inspired by volume-based provenance tracking in Temporal Interaction Networks (TINs), we demonstrate TINs' applicability in succinctly modeling quantified data exchanges between dataflow operators in stream data processing systems and in processing workflows, in general, over time. We classify data into discrete and liquid types, define five temporal provenance query types, and propose a state-based indexing approach. Our vision outlines research directions toward making this new form of temporal attribution a practical tool for large-scale dataflow analytics.
翻译:数据血统(确定数据输出的起源和衍生过程)在多个领域均有应用,包括解释数据库查询结果和审计科学工作流。尽管经过数十年的研究,血统追踪因计算成本高和存储需求大仍面临挑战。在Apache Flink等流式系统中,细粒度的血统图可能随数据量呈超线性增长,带来显著的可扩展性问题。我们定义了时间归因(temporal attribution)——一种适用于特定任务的新型轻量级血统形式,例如定量监测系统组件间随时间变化的依赖关系。时间归因支持无需细粒度元组级依赖元数据的时间聚焦型分析。受时间交互网络(TINs)中基于体积的血统追踪方法启发,我们论证了TINs在流数据处理系统及一般处理工作流中简洁建模数据流算子间随时间变化的量化数据交换的适用性。我们将数据分为离散型和液态型,定义了五种时间血统查询类型,并提出一种基于状态的索引方法。我们的愿景勾勒了使这种新型时间归因成为大规模数据流分析实用工具的研究方向。