The detection of interesting patterns in large high-dimensional datasets is difficult because of their dimensionality and pattern complexity. Therefore, analysts require automated support for the extraction of relevant patterns. In this paper, we present FDive, a visual active learning system that helps to create visually explorable relevance models, assisted by learning a pattern-based similarity. We use a small set of user-provided labels to rank similarity measures, consisting of feature descriptor and distance function combinations, by their ability to distinguish relevant from irrelevant data. Based on the best-ranked similarity measure, the system calculates an interactive Self-Organizing Map-based relevance model, which classifies data according to the cluster affiliation. It also automatically prompts further relevance feedback to improve its accuracy. Uncertain areas, especially near the decision boundaries, are highlighted and can be refined by the user. We evaluate our approach by comparison to state-of-the-art feature selection techniques and demonstrate the usefulness of our approach by a case study classifying electron microscopy images of brain cells. The results show that FDive enhances both the quality and understanding of relevance models and can thus lead to new insights for brain research.
翻译:在高维大数据集中检测有趣模式因数据维度高和模式复杂而困难,因此分析人员需要自动化支持来提取相关模式。本文提出FDive——一个视觉主动学习系统,通过基于模式的相似性学习辅助创建可视觉探索的相关性模型。我们利用少量用户提供的标签对相似性度量(由特征描述符与距离函数组合构成)进行排序,依据其区分相关数据与不相关数据的能力。系统基于排名最优的相似性度量,计算交互式自组织映射相关性模型,依据聚类归属对数据进行分类,并自动提示进一步的相关性反馈以提高准确性。不确定区域(尤其是决策边界附近)会被高亮显示,用户可对其进行精炼。我们通过与前沿特征选择技术对比评估该方法,并通过脑细胞电子显微镜图像分类案例研究证明其实用性。结果表明,FDive能够提升相关性模型的质量与可理解性,从而为脑科学研究带来新见解。