In recent years, video semantic segmentation has made great progress with advanced deep neural networks. However, there still exist two main challenges \ie, information inconsistency and computation cost. To deal with the two difficulties, we propose a novel motion-state alignment framework for video semantic segmentation to keep both motion and state consistency. In the framework, we first construct a motion alignment branch armed with an efficient decoupled transformer to capture dynamic semantics, guaranteeing region-level temporal consistency. Then, a state alignment branch composed of a stage transformer is designed to enrich feature spaces for the current frame to extract static semantics and achieve pixel-level state consistency. Next, by a semantic assignment mechanism, the region descriptor of each semantic category is gained from dynamic semantics and linked with pixel descriptors from static semantics. Benefiting from the alignment of these two kinds of effective information, the proposed method picks up dynamic and static semantics in a targeted way, so that video semantic regions are consistently segmented to obtain precise locations with low computational complexity. Extensive experiments on Cityscapes and CamVid datasets show that the proposed approach outperforms state-of-the-art methods and validates the effectiveness of the motion-state alignment framework.
翻译:近年来,视频语义分割借助先进的深度神经网络取得了重大进展。然而,仍存在两个主要挑战:信息不一致性和计算成本。为应对这两大难题,我们提出了一种新颖的运动状态对齐框架,用于视频语义分割以同时保持运动与状态一致性。在该框架中,我们首先构建一个配备高效解耦Transformer的运动对齐分支,用于捕获动态语义信息,确保区域级时间一致性;随后设计一个由阶段Transformer组成的状态对齐分支,用于丰富当前帧的特征空间以提取静态语义信息,实现像素级状态一致性;接着,通过语义分配机制,从动态语义中获得每个语义类别的区域描述符,并将其与静态语义中的像素描述符相关联。得益于这两种有效信息的对齐,所提方法能够有针对性地提取动态与静态语义,从而实现视频语义区域的连贯分割,以较低的计算复杂度获得精确位置。在Cityscapes和CamVid数据集上的大量实验表明,所提方法优于现有最先进方法,验证了运动状态对齐框架的有效性。