Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these advancements are not solely attributable to novel feature transformation designs; certain benefits also arise from advanced network-level and block-level architectures. This paper aims to identify the real gains of popular convolution and attention operators through a detailed study. We find that the key difference among these feature transformation modules, such as attention or convolution, lies in their spatial feature aggregation approach, known as the "spatial token mixer" (STM). To facilitate an impartial comparison, we introduce a unified architecture to neutralize the impact of divergent network-level and block-level designs. Subsequently, various STMs are integrated into this unified framework for comprehensive comparative analysis. Our experiments on various tasks and an analysis of inductive bias show a significant performance boost due to advanced network-level and block-level designs, but performance differences persist among different STMs. Our detailed analysis also reveals various findings about different STMs, such as effective receptive fields and invariance tests. All models and codes used in this study are publicly available at \url{https://github.com/OpenGVLab/STM-Evaluation}.
翻译:视觉Transformer近年来广受欢迎,催生了具有改进特征和一致性能提升的新型视觉主干网络。然而,这些进展并非完全归因于新颖的特征变换设计;部分优势也源于先进的网络级和块级架构。本文旨在通过详细研究,识别流行卷积与注意力算子的真实增益。我们发现,这些特征变换模块(如注意力或卷积)之间的关键区别在于其空间特征聚合方式,即所谓的“空间令牌混合器”。为促进公平比较,我们引入统一架构以消除不同网络级和块级设计的干扰。随后,我们将多种空间令牌混合器集成到这一统一框架中进行全面比较分析。针对多种任务的实验及归纳偏置分析表明,先进的网络级和块级设计带来了显著性能提升,但不同空间令牌混合器之间仍存在性能差异。我们的详细分析还揭示了关于不同空间令牌混合器的多项发现,例如有效感受野与不变性测试。本研究中使用的所有模型和代码均公开于\url{https://github.com/OpenGVLab/STM-Evaluation}。