Time series remains one of the most challenging modalities in machine learning research. The out-of-distribution (OOD) detection and generalization on time series tend to suffer due to its non-stationary property, i.e., the distribution changes over time. The dynamic distributions inside time series pose great challenges to existing algorithms to identify invariant distributions since they mainly focus on the scenario where the domain information is given as prior knowledge. In this paper, we attempt to exploit subdomains within a whole dataset to counteract issues induced by non-stationary for generalized representation learning. We propose DIVERSIFY, a general framework, for OOD detection and generalization on dynamic distributions of time series. DIVERSIFY takes an iterative process: it first obtains the "worst-case" latent distribution scenario via adversarial training, then reduces the gap between these latent distributions. We implement DIVERSIFY via combining existing OOD detection methods according to either extracted features or outputs of models for detection while we also directly utilize outputs for classification. In addition, theoretical insights illustrate that DIVERSIFY is theoretically supported. Extensive experiments are conducted on seven datasets with different OOD settings across gesture recognition, speech commands recognition, wearable stress and affect detection, and sensor-based human activity recognition. Qualitative and quantitative results demonstrate that DIVERSIFY learns more generalized features and significantly outperforms other baselines.
翻译:时间序列仍然是机器学习研究中最具挑战性的模态之一。由于其非平稳特性(即分布随时间变化),时间序列上的分布外检测与泛化往往面临困难。时间序列内部的动态分布对现有算法识别不变分布构成了重大挑战,因为这些算法主要关注领域信息作为先验知识给出的场景。本文尝试利用整个数据集中的子领域来缓解非平稳性带来的问题,以实现广义表示学习。我们提出DIVERSIFY,这是一个针对时间序列动态分布进行分布外检测与泛化的通用框架。DIVERSIFY采用迭代过程:首先通过对抗训练获取“最坏情况”下的潜在分布场景,然后缩小这些潜在分布之间的差距。我们通过结合现有的分布外检测方法实现DIVERSIFY,这些方法基于提取的特征或模型输出进行检测,同时我们也直接利用输出进行分类。此外,理论分析表明DIVERSIFY具有理论基础。在七个不同分布外设置的数据集上进行了广泛实验,涵盖手势识别、语音命令识别、可穿戴压力与情绪检测以及基于传感器的人体活动识别。定性和定量结果表明,DIVERSIFY学习了更泛化的特征,并在性能上显著优于其他基线方法。