Visualizations for scattered data are used to make users understand certain attributes of their data by solving different tasks, e.g. correlation estimation, outlier detection, cluster separation. In this paper, we focus on the later task, and develop a technique that is aligned to human perception, that can be used to understand how human subjects perceive clusterings in scattered data and possibly optimize for better understanding. Cluster separation in scatterplots is a task that is typically tackled by widely used clustering techniques, such as for instance k-means or DBSCAN. However, as these algorithms are based on non-perceptual metrics, we can show in our experiments, that their output do not reflect human cluster perception. We propose a learning strategy which directly operates on scattered data. To learn perceptual cluster separation on this data, we crowdsourced a large scale dataset, consisting of 7,320 point-wise cluster affiliations for bivariate data, which has been labeled by 384 human crowd workers. Based on this data, we were able to train ClusterNet, a point-based deep learning model, trained to reflect human perception of cluster separability. In order to train ClusterNet on human annotated data, we use a PointNet++ architecture enabling inference on point clouds directly. In this work, we provide details on how we collected our dataset, report statistics of the resulting annotations, and investigate perceptual agreement of cluster separation for real-world data. We further report the training and evaluation protocol of ClusterNet and introduce a novel metric, that measures the accuracy between a clustering technique and a group of human annotators. Finally, we compare our approach against existing state-of-the-art clustering techniques and can show, that ClusterNet is able to generalize to unseen and out of scope data.
翻译:摘要:散乱数据的可视化旨在通过解决不同任务(如相关性估计、异常值检测、簇分离)帮助用户理解其数据的特定属性。本文聚焦于簇分离任务,开发了一种符合人类感知的技术,可用于理解人类受试者如何感知散乱数据中的聚类现象,并可能为优化理解提供支持。散点图中的簇分离通常由广泛应用的聚类技术处理,例如k-means或DBSCAN。然而,这些算法基于非感知性度量,我们在实验中证明其输出无法反映人类对聚类的感知。为此,我们提出了一种直接作用于散乱数据的学习策略。为学习该数据上的感知性簇分离,我们通过众包构建了一个大规模数据集,包含7,320个双变量数据的逐点簇隶属标注,这些标注由384名众包工作者完成。基于该数据,我们训练了ClusterNet——一种基于点的深度学习模型,旨在反映人类对簇可分性的感知。为在人工标注数据上训练ClusterNet,我们采用PointNet++架构,使其能够直接对点云进行推理。本文详细介绍了数据集的收集过程、标注结果的统计数据,并探究了真实世界数据中簇分离的感知一致性。我们进一步报告了ClusterNet的训练与评估协议,并提出了一种新型度量方法,用于衡量聚类技术与人类标注者群体之间的一致性。最后,我们将所提方法与现有最先进的聚类技术进行比较,结果表明ClusterNet能够泛化至未见及域外数据。