Multi-instance learning (MIL) is a widely-applied technique in practical applications that involve complex data structures. MIL can be broadly categorized into two types: traditional methods and those based on deep learning. These approaches have yielded significant results, especially with regards to their problem-solving strategies and experimental validation, providing valuable insights for researchers in the MIL field. However, a considerable amount of knowledge is often trapped within the algorithm, leading to subsequent MIL algorithms that solely rely on the model's data fitting to predict unlabeled samples. This results in a significant loss of knowledge and impedes the development of more intelligent models. In this paper, we propose a novel data-driven knowledge fusion for deep multi-instance learning (DKMIL) algorithm. DKMIL adopts a completely different idea from existing deep MIL methods by analyzing the decision-making of key samples in the data set (referred to as the data-driven) and using the knowledge fusion module designed to extract valuable information from these samples to assist the model's training. In other words, this module serves as a new interface between data and the model, providing strong scalability and enabling the use of prior knowledge from existing algorithms to enhance the learning ability of the model. Furthermore, to adapt the downstream modules of the model to more knowledge-enriched features extracted from the data-driven knowledge fusion module, we propose a two-level attention module that gradually learns shallow- and deep-level features of the samples to achieve more effective classification. We will prove the scalability of the knowledge fusion module while also verifying the efficacy of the proposed architecture by conducting experiments on 38 data sets across 6 categories.
翻译:多实例学习(MIL)是一种广泛应用于涉及复杂数据结构的实际问题中的技术。MIL大致可分为两类:传统方法和基于深度学习的方法。这些方法已取得显著成果,尤其是在问题解决策略和实验验证方面,为MIL领域的研究人员提供了宝贵见解。然而,大量知识常常被困在算法内部,导致后续MIL算法仅依赖模型的数据拟合来预测未标记样本。这造成了显著的知识损失,并阻碍了更智能模型的发展。本文提出了一种新颖的数据驱动深度多实例学习知识融合(DKMIL)算法。DKMIL采用与现有深度MIL方法完全不同的思路,通过分析数据集中关键样本的决策(称为数据驱动),并利用设计的知识融合模块从这些样本中提取有价值信息以辅助模型训练。换言之,该模块作为数据与模型之间的新接口,具备强大的可扩展性,并能利用现有算法中的先验知识来增强模型的学习能力。此外,为使模型的下游模块适应从数据驱动知识融合模块中提取的更富含知识的特征,我们提出了一种两级注意力模块,该模块逐步学习样本的浅层和深层特征,以实现更有效的分类。我们将证明知识融合模块的可扩展性,同时通过在6大类别共38个数据集上开展实验来验证所提架构的有效性。