A data analysis pipeline is a structured sequence of steps that transforms raw data into meaningful insights by integrating multiple analysis algorithms. In many practical applications, analytical findings are obtained only after data pass through several data-dependent procedures within such pipelines. In this study, we address the problem of quantifying the statistical reliability of results produced by data analysis pipelines. As a proof of concept, we focus on clustering pipelines that identify cluster structures from complex and heterogeneous data through procedures such as outlier detection, feature selection, and clustering. We propose a novel statistical testing framework to assess the significance of clustering results obtained through these pipelines. Our framework, based on selective inference, enables the systematic construction of valid statistical tests for clustering pipelines composed of predefined components. We prove that the proposed test controls the type I error rate at any nominal level and demonstrate its validity and effectiveness through experiments on synthetic and real datasets.
翻译:数据分析流程是一个将原始数据通过整合多种分析算法转化为有意义洞察的结构化步骤序列。在许多实际应用中,分析结果仅在数据经过此类流程中的多个数据依赖步骤后才能获得。本研究针对量化数据分析流程产出结果的统计可靠性问题展开探讨。作为概念验证,我们聚焦于通过异常值检测、特征选择和聚类等步骤从复杂异构数据中识别聚类结构的聚类流程。我们提出了一种新颖的统计检验框架,用于评估通过此类流程获得的聚类结果的显著性。该框架基于选择性推断,能够系统性地为包含预定义组件的聚类流程构建有效的统计检验。我们证明了所提检验方法能在任意名义水平上控制第一类错误率,并通过合成数据集与真实数据集的实验验证了其有效性和实用性。