Recently, state space models have demonstrated efficient video segmentation through linear-complexity state space compression. However, Video Semantic Segmentation (VSS) requires pixel-level spatiotemporal modeling capabilities to maintain temporal consistency in segmentation of semantic objects. While state space models can preserve common semantic information during state space compression, the fixed-size state space inevitably forgets specific information, which limits the models' capability for pixel-level segmentation. To tackle the above issue, we proposed a Refining Specifics State Space Model approach (RS-SSM) for video semantic segmentation, which performs complementary refining of forgotten spatiotemporal specifics. Specifically, a Channel-wise Amplitude Perceptron (CwAP) is designed to extract and align the distribution characteristics of specific information in the state space. Besides, a Forgetting Gate Information Refiner (FGIR) is proposed to adaptively invert and refine the forgetting gate matrix in the state space model based on the specific information distribution. Consequently, our RS-SSM leverages the inverted forgetting gate to complementarily refine the specific information forgotten during state space compression, thereby enhancing the model's capability for spatiotemporal pixel-level segmentation. Extensive experiments on four VSS benchmarks demonstrate that our RS-SSM achieves state-of-the-art performance while maintaining high computational efficiency. The code is available at https://github.com/zhoujiahuan1991/CVPR2026-RS-SSM.


翻译:近期,状态空间模型通过线性复杂度的状态空间压缩展示了高效的视频分割能力。然而,视频语义分割(VSS)需要像素级的时空建模能力,以维持语义对象分割的时间一致性。尽管状态空间模型在状态空间压缩过程中能够保留公共语义信息,但固定大小的状态空间不可避免地会遗忘特定信息,这限制了模型在像素级分割中的能力。为解决上述问题,我们提出了一种用于视频语义分割的精炼特定细节状态空间模型方法(RS-SSM),该方法对遗忘的时空特定细节进行互补性精炼。具体而言,我们设计了一个通道级幅度感知器(CwAP),用于提取和对齐状态空间中特定信息的分布特征。此外,我们还提出了一种遗忘门信息精炼器(FGIR),基于特定信息分布自适应地反转和精炼状态空间模型中的遗忘门矩阵。由此,我们的RS-SSM利用反转的遗忘门对状态空间压缩过程中遗忘的特定细节进行互补性精炼,从而增强了模型在时空像素级分割中的能力。在四个VSS基准上的大量实验表明,我们的RS-SSM在保持高计算效率的同时,实现了最先进的性能。代码已在https://github.com/zhoujiahuan1991/CVPR2026-RS-SSM公开。

0
下载
关闭预览

相关内容

视频大模型中视觉上下文表示的scaling law
专知会员服务
24+阅读 · 2024年10月21日
非Transformer不可?最新《状态空间模型(SSM)》综述
专知会员服务
75+阅读 · 2024年4月16日
基于深度学习的实时语义分割综述
专知会员服务
33+阅读 · 2023年11月27日
专知会员服务
23+阅读 · 2021年7月5日
AAAI2021 | DTGRM:具有自监督时间关系建模的动作分割
专知会员服务
15+阅读 · 2020年12月29日
综述 | 语义分割经典网络及轻量化模型盘点
计算机视觉life
54+阅读 · 2019年7月23日
DL | 语义分割综述
机器学习算法与Python学习
58+阅读 · 2019年3月13日
超像素、语义分割、实例分割、全景分割 傻傻分不清?
计算机视觉life
19+阅读 · 2018年11月27日
语义分割中的深度学习方法全解:从FCN、SegNet到DeepLab
炼数成金订阅号
26+阅读 · 2017年7月10日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
VIP会员
最新内容
印度精确打击与指挥架构的断层
专知会员服务
4+阅读 · 7月20日
美空军AI完成F-16战斗机自主空战历史性试飞
专知会员服务
5+阅读 · 7月20日
深入Project Maven:为何人工智能在战场上依然失灵
锻造未来士兵:外骨骼、基因工程与赛博格
专知会员服务
7+阅读 · 7月19日
《无人机蜂群通信技术研究》50页
专知会员服务
8+阅读 · 7月19日
相关VIP内容
视频大模型中视觉上下文表示的scaling law
专知会员服务
24+阅读 · 2024年10月21日
非Transformer不可?最新《状态空间模型(SSM)》综述
专知会员服务
75+阅读 · 2024年4月16日
基于深度学习的实时语义分割综述
专知会员服务
33+阅读 · 2023年11月27日
专知会员服务
23+阅读 · 2021年7月5日
AAAI2021 | DTGRM:具有自监督时间关系建模的动作分割
专知会员服务
15+阅读 · 2020年12月29日
相关基金
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
5+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员