In the biomedical domain, taxonomies organize the acquisition modalities of scientific images in hierarchical structures. Such taxonomies leverage large sets of correct image labels and provide essential information about the importance of a scientific publication, which could then be used in biocuration tasks. However, the hierarchical nature of the labels, the overhead of processing images, the absence or incompleteness of labeled data, and the expertise required to label this type of data impede the creation of useful datasets for biocuration. From a multi-year collaboration with biocurators and text-mining researchers, we derive an iterative visual analytics and active learning strategy to address these challenges. We implement this strategy in a system called BI-LAVA Biocuration with Hierarchical Image Labeling through Active Learning and Visual Analysis. BI-LAVA leverages a small set of image labels, a hierarchical set of image classifiers, and active learning to help model builders deal with incomplete ground-truth labels, target a hierarchical taxonomy of image modalities, and classify a large pool of unlabeled images. BI-LAVA's front end uses custom encodings to represent data distributions, taxonomies, image projections, and neighborhoods of image thumbnails, which help model builders explore an unfamiliar image dataset and taxonomy and correct and generate labels. An evaluation with machine learning practitioners shows that our mixed human-machine approach successfully supports domain experts in understanding the characteristics of classes within the taxonomy, as well as validating and improving data quality in labeled and unlabeled collections.
翻译:在生物医学领域,分类法将科学图像的获取模态组织成层级结构。此类分类法利用大量正确图像标签,提供关于科学出版物重要性的关键信息,这些信息可用于生物管理任务。然而,标签的层级特性、图像处理的额外开销、标注数据的缺失或不完整性,以及对此类数据进行标注所需的专业知识,阻碍了生物管理有用数据集的创建。基于与生物管理员和文本挖掘研究人员历时多年的合作,我们提出了一种迭代式视觉分析与主动学习策略来应对这些挑战。我们通过名为BI-LAVA的系统(Biocuration with Hierarchical Image Labeling through Active Learning and Visual Analysis,即基于主动学习与视觉分析的分级图像标注生物管理)实现该策略。BI-LAVA利用少量图像标签、一组层级式图像分类器及主动学习,帮助模型构建者处理不完整的地面真值标签,面向图像模态的层级分类法进行分类,并对大量未标注图像进行归类。BI-LAVA的前端采用自定义编码来表示数据分布、分类法、图像投影及图像缩略图的邻域结构,助力模型构建者探索不熟悉的图像数据集和分类法,修正并生成标签。通过机器学习从业者的评估表明,我们的人机混合方法成功支持领域专家理解分类法中各类别的特征,验证并改善标注与未标注集合的数据质量。