Real-world instructional videos are long, noisy, and often contain extended background segments, repeated actions, and execution variability that do not correspond to meaningful procedural steps. We propose **REMAP**, an unsupervised framework for procedure learning based on *Regularized Fused Partial Gromov-Wasserstein Optimal Transport*. REMAP relaxes balanced transport constraints, allowing non-informative or redundant frames to remain unmatched through partial transport. The formulation jointly models semantic similarity and temporal structure, while incorporating Laplacian-based smoothness and structural regularization to prevent degenerate alignments and reduce background interference. We evaluate REMAP on large-scale egocentric and third-person benchmarks. The method consistently outperforms state-of-the-art approaches, achieving up to **11.6\% (+4.45pp)** F1 and **19.6\% (+4.73pp)** IoU improvements on EgoProceL, and an average **41\% (+17.15pp)** F1 gain on ProceL and CrossTask. These results highlight the importance of partial alignment in handling real-world procedural variability and demonstrate that REMAP provides a robust and scalable approach for instructional video understanding.
翻译:摘要:真实教学视频通常冗长、嘈杂,且常包含无关背景片段、重复动作以及执行变异性,这些内容并不对应有意义的程序步骤。我们提出**REMAP**——一种基于正则化融合部分Gromov-Wasserstein最优传输的无监督程序学习框架。REMAP通过放宽平衡传输约束,利用部分传输机制使非信息性帧或冗余帧保持未匹配状态。该公式联合建模语义相似性与时序结构,同时引入基于拉普拉斯平滑的结构正则化,以防止退化对齐并减少背景干扰。我们在大规模第一人称与第三人称基准数据集上评估了REMAP。该方法持续优于现有最先进技术,在EgoProceL数据集上F1分数提升最高达**11.6%(+4.45个百分点)**,IoU提升**19.6%(+4.73个百分点)**;在ProceL与CrossTask数据集上平均F1分数提升**41%(+17.15个百分点)**。这些结果凸显了部分对齐在处理真实世界程序变异性中的关键作用,并证明REMAP为教学视频理解提供了鲁棒且可扩展的解决方案。