Whether by processing videos with fixed resolution from start to end or incorporating pooling and down-scaling strategies, existing video transformers process the whole video content throughout the network without specially handling the large portions of redundant information. In this paper, we present a Supertoken Video Transformer (SVT) that incorporates a Semantic Pooling Module (SPM) to aggregate latent representations along the depth of visual transformer based on their semantics, and thus, reduces redundancy inherent in video inputs.~Qualitative results show that our method can effectively reduce redundancy by merging latent representations with similar semantics and thus increase the proportion of salient information for downstream tasks.~Quantitatively, our method improves the performance of both ViT and MViT while requiring significantly less computations on the Kinectics and Something-Something-V2 benchmarks.~More specifically, with our SPM, we improve the accuracy of MAE-pretrained ViT-B and ViT-L by 1.5% with 33% less GFLOPs and by 0.2% with 55% less FLOPs, respectively, on the Kinectics-400 benchmark, and improve the accuracy of MViTv2-B by 0.2% and 0.3% with 22% less GFLOPs on Kinectics-400 and Something-Something-V2, respectively.
翻译:现有视频Transformer或从始至终以固定分辨率处理视频,或采用池化与下采样策略,但均未专门处理视频中大量存在的冗余信息,而是在整个网络中完整处理视频内容。本文提出超令牌视频Transformer(SVT),通过引入语义池化模块(SPM),基于语义信息沿视觉Transformer深度方向聚合潜在表征,从而降低视频输入固有的冗余性。定性结果表明,该方法能有效合并具有相似语义的潜在表征以降低冗余性,进而提升下游任务中显著信息的占比。定量实验显示,在Kinetics和Something-Something-V2基准测试中,该方法在显著降低计算量的同时,提升了ViT和MViT的性能。具体而言,在Kinetics-400基准上,所提SPM使MAE预训练ViT-B和ViT-L的准确率分别提升1.5%(计算量降低33%GFLOPs)和0.2%(计算量降低55%FLOPs);在Kinetics-400和Something-Something-V2基准上,使MViTv2-B的准确率分别提升0.2%和0.3%(计算量均降低22%GFLOPs)。