For artificial intelligence, high-utility sequential rule mining (HUSRM) is a knowledge discovery method that can reveal the associations between events in the sequences. Recently, abundant methods have been proposed to discover high-utility sequence rules. However, the existing methods are all related to point-based sequences. Interval events that persist for some time are common. Traditional interval-event sequence knowledge discovery tasks mainly focus on pattern discovery, but patterns cannot reveal the correlation between interval events well. Moreover, the existing HUSRM algorithms cannot be directly applied to interval-event sequences since the relation in interval-event sequences is much more intricate than those in point-based sequences. In this work, we propose a utility-driven interval rule mining (UIRMiner) algorithm that can extract all utility-driven interval rules (UIRs) from the interval-event sequence database to solve the problem. In UIRMiner, we first introduce a numeric encoding relation representation, which can save much time on relation computation and storage on relation representation. Furthermore, to shrink the search space, we also propose a complement pruning strategy, which incorporates the utility upper bound with the relation. Finally, plentiful experiments implemented on both real-world and synthetic datasets verify that UIRMiner is an effective and efficient algorithm.
翻译:对于人工智能而言,高效用序列规则挖掘(HUSRM)是一种能够揭示序列中事件之间关联的知识发现方法。近年来,已有大量方法被提出用于发现高效用序列规则。然而,现有方法均与基于点的序列相关。持续一段时间的区间事件是普遍存在的。传统的区间事件序列知识发现任务主要侧重于模式发现,但模式无法很好揭示区间事件之间的关联性。此外,现有HUSRM算法无法直接应用于区间事件序列,因为区间事件序列中的关系比基于点的序列中的关系复杂得多。本文提出了一种效用驱动的区间规则挖掘(UIRMiner)算法,该算法能够从区间事件序列数据库中提取所有效用驱动的区间规则(UIRs)以解决上述问题。在UIRMiner中,我们首先引入了一种数值编码的关系表示方法,该方法可在关系计算和关系表示存储方面节省大量时间。此外,为了缩小搜索空间,我们还提出了一种结合效用上界与关系的互补剪枝策略。最后,在真实数据集和合成数据集上进行的丰富实验验证了UIRMiner是一种高效且有效的算法。