Real-world temporal data often consists of multiple signal types recorded at irregular, asynchronous intervals. For instance, in the medical domain, different types of blood tests can be measured at different times and frequencies, resulting in fragmented and unevenly scattered temporal data. Similar issues of irregular sampling occur in other domains, such as the monitoring of large systems using event log files. Effectively learning from such data requires handling sets of temporal sparse and heterogeneous signals. In this work, we propose Super Mixing Additive Networks (SuperMAN), a novel and interpretable-by-design framework for learning directly from such heterogeneous signals, by modeling them as sets of implicit graphs. SuperMAN provides diverse interpretability capabilities, including node-level, graph-level, and subset-level importance, and enables practitioners to trade finer-grained interpretability for greater expressivity when domain priors are available. SuperMAN achieves state-of-the-art performance in real-world high-stakes tasks, including predicting Crohn's disease onset and hospital length of stay from routine blood test measurements and detecting fake news. Furthermore, we demonstrate how SuperMAN's interpretability properties assist in revealing disease development phase transitions and provide crucial insights in the healthcare domain.
翻译:现实世界中的时序数据通常包含多种以不规则、异步间隔记录的信号类型。例如,在医疗领域中,不同类型的血液检测可在不同时间和频率下进行测量,从而产生碎片化且分布不均的时序数据。类似的非规则采样问题也出现在其他领域,例如使用事件日志文件对大型系统进行监测。要从这类数据中有效学习,需要处理时序稀疏且异构的信号集合。本文提出超级混合可加网络(SuperMAN),这是一种新颖且设计上具有可解释性的框架,通过将异构信号建模为隐式图集合,可直接从此类信号中进行学习。SuperMAN提供多样化的可解释能力,包括节点级、图级和子集级重要性分析,并允许从业者在具备领域先验知识时,通过调整更细粒度的可解释性来获得更强的表达能力。SuperMAN在现实世界的高风险任务中实现了最先进的性能,包括基于常规血液检测指标预测克罗恩病发病与住院时长,以及虚假新闻检测。此外,我们展示了SuperMAN的可解释性特性如何帮助揭示疾病发展的阶段转变,并为医疗健康领域提供关键洞见。