Machine learning holds tremendous promise for transforming the fundamental practice of scientific discovery by virtue of its data-driven nature. With the ever-increasing stream of research data collection, it would be appealing to autonomously explore patterns and insights from observational data for discovering novel classes of phenotypes and concepts. However, in the biomedical domain, there are several challenges inherently presented in the cumulated data which hamper the progress of novel class discovery. The non-i.i.d. data distribution accompanied by the severe imbalance among different groups of classes essentially leads to ambiguous and biased semantic representations. In this work, we present a geometry-constrained probabilistic modeling treatment to resolve the identified issues. First, we propose to parameterize the approximated posterior of instance embedding as a marginal von MisesFisher distribution to account for the interference of distributional latent bias. Then, we incorporate a suite of critical geometric properties to impose proper constraints on the layout of constructed embedding space, which in turn minimizes the uncontrollable risk for unknown class learning and structuring. Furthermore, a spectral graph-theoretic method is devised to estimate the number of potential novel classes. It inherits two intriguing merits compared to existent approaches, namely high computational efficiency and flexibility for taxonomy-adaptive estimation. Extensive experiments across various biomedical scenarios substantiate the effectiveness and general applicability of our method.
翻译:机器学习凭借其数据驱动的特性,在改变科学发现这一基础实践方面展现出巨大潜力。随着研究数据收集的持续增长,自主从观测数据中探索模式与洞见,从而发现新型表型与概念类别,具有诱人前景。然而在生物医学领域,累积数据中固有的若干挑战阻碍了新型类别发现的进展。非独立同分布的数据分布伴随不同类别群体间的严重不平衡,本质上导致语义表征模糊且存在偏差。本研究提出一种几何约束概率建模方法以解决上述问题。首先,我们提出将实例嵌入的近似后验参数化为边缘冯·米塞斯-费希尔分布,以处理分布潜在偏差的干扰。随后,我们纳入一组关键几何性质,对构建的嵌入空间布局施加适当约束,从而最小化未知类别学习与结构化过程中的不可控风险。此外,我们设计了一种基于谱图理论的方法来估计潜在新型类别的数量。与现有方法相比,该方法具有两个显著优势:计算效率高且支持分类学自适应估计。跨多种生物医学场景的大量实验验证了我们方法的有效性与通用适用性。