Outliers are known to significantly distort the results of many commonly used clustering methods, often leading to unreliable partitions. To address this issue, several robust clustering approaches have been developed that not only reduce their influence but also facilitate the detection of meaningful outliers. This presentation focuses on robust clustering methods based on trimming, especially TCLUST, which extends the type of trimming used by MCD in one-population problems to the more general case of multiple and unknown clusters. While TCLUST performs well on low-dimensional data, it struggles with high-dimensional datasets due to the complexity of estimating a large number of parameters. The Robust Linear Grouping (RLG) method offers an alternative by assuming clusters lie near lower-dimensional subspaces, thereby combining clustering with dimensionality reduction. However, RLG has limitations when subspaces intersect and assumes overly simplistic isotropic orthogonal errors. A robust clustering method extending TCLUST will be presented, building on the High Dimensional Data Clustering (HDDC) approach by incorporating trimming and eigenvalue constraints. This new approach, called tHHDC, combines TCLUST and RLG, requiring careful modification and integration of both methodologies within that HDDC framework. A study of the theoretical properties of this approach, together with a feasible algorithm for its implementation, will be presented. The interest of the proposed methodology, along with the issue of selecting input parameters, will be illustrated through a simulation study and a real-data example.
翻译:已知异常值会显著扭曲许多常用聚类方法的结果,常导致不可靠的分割。为解决这一问题,研究人员开发了若干稳健聚类方法,这些方法不仅能降低异常值的影响,还能促进有意义异常值的检测。本报告聚焦于基于修整的稳健聚类方法,特别是TCLUST,它将单总体问题中MCD所用的修整类型推广至多个未知聚类的更一般情形。尽管TCLUST在低维数据上表现良好,但由于估计大量参数的复杂性,它在高维数据集上存在困难。稳健线性分组方法通过假设聚类位于较低维子空间附近,从而将聚类与降维相结合,提供了一种替代方案。然而,RLG在子空间相交时存在局限性,并假定了过于简单的各向同性正交误差。本文将提出一种扩展TCLUST的稳健聚类方法,该方法借鉴高维数据聚类思路,通过引入修整和特征值约束进行改进。这一新方法称为tHHDC,它结合了TCLUST与RLG,需要在HDDC框架内对这两种方法进行仔细修改与整合。本文将对该方法的理论性质进行研究,并给出可行的实现算法。通过模拟研究和实际数据案例,将展示所提方法的有效性以及输入参数选择问题的重要性。