This work uses visual knowledge discovery in parallel coordinates to advance methods of interpretable machine learning. The graphic data representation in parallel coordinates made the concepts of hypercubes and hyperblocks (HBs) simple to understand for end users. It is suggested to use mixed and pure hyperblocks in the proposed data classifier algorithm Hyper. It is shown that Hyper models generalize decision trees. The algorithm is presented in several settings and options to discover interactively or automatically overlapping or non-overlapping hyperblocks. Additionally, the use of hyperblocks in conjunction with language descriptions of visual patterns is demonstrated. The benchmark data from the UCI ML repository were used to evaluate the Hyper algorithm. It enabled the discovery of mixed and pure HBs evaluated using 10-fold cross validation. Connections among hyperblocks, dimension reduction and visualization have been established. The capability of end users to find and observe hyperblocks, as well as the ability of side-by-side visualizations to make patterns evident, are among major advantages ofhyperblock technology and the Hyper algorithm. A new method to visualize incomplete n-D data with missing values is proposed, while the traditional parallel coordinates do not support it. The ability of HBs to better prevent both overgeneralization and overfitting of data over decision trees is demonstrated as another benefit of the hyperblocks. The features of VisCanvas 2.0 software tool that implements Hyper technology are presented.
翻译:本文利用平行坐标中的可视化知识发现,推进了可解释机器学习方法的研究。平行坐标中的图形数据表示使得超立方体和超块(HBs)的概念对最终用户易于理解。建议在所提出的数据分类算法Hyper中使用混合超块和纯超块。研究表明,Hyper模型泛化了决策树。该算法以多种设置和选项呈现,用于交互式或自动发现重叠或非重叠超块。此外,还展示了超块与视觉模式语言描述的结合使用。使用UCI机器学习库中的基准数据对Hyper算法进行了评估。该算法能够发现通过10折交叉验证评估的混合与纯超块。超块、降维和可视化之间的联系已被建立。最终用户发现和观察超块的能力,以及并排可视化使模式显现的能力,是超块技术与Hyper算法的主要优势之一。本文提出了一种可视化包含缺失值的不完整n维数据的新方法,而传统平行坐标不支持此功能。作为超块的另一优势,本文证明了超块比决策树更能有效防止数据的过度泛化和过拟合。本文还介绍了实现Hyper技术的VisCanvas 2.0软件工具的功能。