Weakly-supervised temporal action localization aims to localize action instances in untrimmed videos with only video-level supervision. We witness that different actions record common phases, e.g., the run-up in the HighJump and LongJump. These different actions are defined as conjoint actions, whose rest parts are definite phases, e.g., leaping over the bar in a HighJump. Compared with the common phases, the definite phases are more easily localized in existing researches. Most of them formulate this task as a Multiple Instance Learning paradigm, in which the common phases are tended to be confused with the background, and affect the localization completeness of the conjoint actions. To tackle this challenge, we propose a Joint of Common and Definite phases Network (JCDNet) by improving feature discriminability of the conjoint actions. Specifically, we design a Class-Aware Discriminative module to enhance the contribution of the common phases in classification by the guidance of the coarse definite-phase features. Besides, we introduce a temporal attention module to learn robust action-ness scores via modeling temporal dependencies, distinguishing the common phases from the background. Extensive experiments on three datasets (THUMOS14, ActivityNetv1.2, and a conjoint-action subset) demonstrate that JCDNet achieves competitive performance against the state-of-the-art methods. Keywords: weakly-supervised learning, temporal action localization, conjoint action
翻译:弱监督时序动作定位旨在仅利用视频级监督信息在未修剪视频中定位动作实例。我们观察到不同动作记录有公共阶段,例如跳高和跳远中的助跑。这些不同动作被定义为联合动作,其其余部分为确定阶段,例如跳高中跨越横杆的动作。与公共阶段相比,现有研究中确定阶段更容易被定位。多数方法将这一任务建模为多实例学习范式,其中公共阶段易与背景混淆,并影响联合动作的定位完整性。为应对这一挑战,我们提出联合公共阶段与确定阶段网络(JCDNet),通过提升联合动作的特征判别性来解决问题。具体而言,我们设计类感知判别模块,利用粗粒度确定阶段特征引导增强公共阶段在分类中的贡献。此外,引入时序注意力模块,通过建模时序依赖关系学习鲁棒的动作性分数,从而区分公共阶段与背景。在三个数据集(THUMOS14、ActivityNetv1.2及联合动作子集)上的大量实验表明,JCDNet相较于现有最优方法取得了具有竞争力的性能。关键词:弱监督学习,时序动作定位,联合动作