We establish Multilayer Correlation Clustering, a novel generalization of Correlation Clustering to the multilayer setting. In this model, we are given a series of inputs of Correlation Clustering (called layers) over the common set $V$ of $n$ elements. The goal is to find a clustering of $V$ that minimizes the $\ell_p$-norm ($p\geq 1$) of the multilayer-disagreements vector, which is defined as the vector (with dimension equal to the number of layers), each element of which represents the disagreements of the clustering on the corresponding layer. For this generalization, we first design an $O(L\log n)$-approximation algorithm, where $L$ is the number of layers. We then study an important special case of our problem, namely the problem with the so-called probability constraint. For this case, we first give an $(α+2)$-approximation algorithm, where $α$ is any possible approximation ratio for the single-layer counterpart. Furthermore, we design a $4$-approximation algorithm, which improves the above approximation ratio of $α+2=4.5$ for the general probability-constraint case. Computational experiments using real-world datasets support our theoretical findings and demonstrate the practical effectiveness of our proposed algorithms.
翻译:我们建立了多层相关聚类(Multilayer Correlation Clustering),这是对相关聚类(Correlation Clustering)在多层级设置下的新推广。在该模型中,我们获得一系列基于公共元素集$V$(包含$n$个元素)的相关聚类输入(称为层)。目标是在$V$上找到一个聚类,使得多层不一致向量的$\ell_p$范数($p\geq 1$)最小化。该向量(其维度等于层数)的每个元素表示聚类在对应层上的不一致程度。针对这一推广,我们首先设计了一种$O(L\log n)$近似算法,其中$L$为层数。随后,我们研究该问题的一个重要特例——即具有所谓概率约束的问题。对于该特例,我们首先给出一种$(\alpha+2)$近似算法,其中$\alpha$为单层对应问题的任意可能近似比。此外,我们设计了一种$4$近似算法,将通用概率约束情况下的上述近似比$\alpha+2=4.5$进一步改进。基于真实数据集的实验验证了我们的理论发现,并展示了所提算法的实际有效性。