Transformers are widely used for solving tasks in natural language processing, computer vision, speech, and music domains. In this paper, we talk about the efficiency of transformers in terms of memory (the number of parameters), computation cost (number of floating points operations), and performance of models, including accuracy, the robustness of the model, and fair \& bias-free features. We mainly discuss the vision transformer for the image classification task. Our contribution is to introduce an efficient 360 framework, which includes various aspects of the vision transformer, to make it more efficient for industrial applications. By considering those applications, we categorize them into multiple dimensions such as privacy, robustness, transparency, fairness, inclusiveness, continual learning, probabilistic models, approximation, computational complexity, and spectral complexity. We compare various vision transformer models based on their performance, the number of parameters, and the number of floating point operations (FLOPs) on multiple datasets.
翻译:Transformer被广泛用于解决自然语言处理、计算机视觉、语音和音乐领域的任务。本文从内存(参数数量)、计算成本(浮点运算次数)以及模型性能(包括准确性、鲁棒性以及公平与无偏特征)三个维度探讨Transformer的效率问题。我们主要针对图像分类任务中的视觉Transformer进行讨论。本文的贡献在于引入一个高效的360框架,该框架涵盖视觉Transformer的多个方面,以提升其在工业应用中的效率。基于这些应用,我们将其归为多个维度,包括隐私性、鲁棒性、透明性、公平性、包容性、持续学习、概率模型、近似算法、计算复杂度以及谱复杂度。我们在多个数据集上对比了不同视觉Transformer模型的性能、参数数量及浮点运算次数(FLOPs)。