Selecting the number of clusters remains a fundamental challenge in unsupervised learning. Existing criteria typically target a single ``optimal'' partition, often overlooking statistically meaningful structure present at multiple resolutions. We introduce ElbowSig, a framework that formalizes the heuristic ``elbow'' method as a rigorous inferential problem. Our approach centers on a normalized discrete curvature statistic derived from the cluster heterogeneity sequence, which is evaluated against a null distribution of unstructured data. We derive the asymptotic properties of this null statistic in both large-sample and high-dimensional regimes, characterizing its baseline behavior and stochastic variability. As an algorithm-agnostic procedure, ElbowSig requires only the heterogeneity sequence and is compatible with a wide range of clustering methods, including hard, fuzzy, and model-based clustering. Extensive experiments on synthetic and empirical datasets demonstrate that the method maintains appropriate Type-I error control while providing the power to resolve multiscale organizational structures that are typically obscured by single-resolution selection criteria.
翻译:在无监督学习中,选择聚类数量仍然是一个根本性挑战。现有准则通常针对单一的“最优”划分,往往忽略了存在于多个分辨率下的具有统计意义的结构。我们提出了ElbowSig框架,将启发式的“肘部”方法形式化为一个严格的推断问题。我们的方法核心在于从聚类异质性序列导出的归一化离散曲率统计量,该统计量通过与非结构化数据的零分布进行比较来评估。我们推导了该零统计量在大样本和高维情况下的渐近性质,刻画了其基线行为和随机变异性。作为一种与算法无关的流程,ElbowSig仅需要异质性序列,并且兼容包括硬聚类、模糊聚类和基于模型的聚类在内的多种聚类方法。在合成数据集和实证数据集上的大量实验表明,该方法在保持适当的I类错误控制的同时,能够有效解析通常被单分辨率选择准则所掩盖的多尺度组织结构。