In this paper, we propose a novel fully unsupervised framework that learns action representations suitable for the action segmentation task from the single input video itself, without requiring any training data. Our method is a deep metric learning approach rooted in a shallow network with a triplet loss operating on similarity distributions and a novel triplet selection strategy that effectively models temporal and semantic priors to discover actions in the new representational space. Under these circumstances, we successfully recover temporal boundaries in the learned action representations with higher quality compared with existing unsupervised approaches. The proposed method is evaluated on two widely used benchmark datasets for the action segmentation task and it achieves competitive performance by applying a generic clustering algorithm on the learned representations.
翻译:本文提出了一种全新的全无监督框架,该框架无需任何训练数据,仅从单一输入视频中学习适用于动作分割任务的动作表示。我们的方法是一种深度度量学习方法,其核心是一个基于相似度分布上三元组损失的浅层网络,并采用了一种新颖的三元组选择策略,该策略能有效建模时序和语义先验,从而在新型表示空间中发现动作。在此条件下,与现有无监督方法相比,我们成功地从学到的动作表示中恢复出更高质量的时序边界。该方法在两个广泛使用的动作分割基准数据集上进行了评估,通过对学到的表示应用通用聚类算法,取得了具有竞争力的性能。