Learning from a limited amount of data, namely Few-Shot Learning, stands out as a challenging computer vision task. Several works exploit semantics and design complicated semantic fusion mechanisms to compensate for rare representative features within restricted data. However, relying on naive semantics such as class names introduces biases due to their brevity, while acquiring extensive semantics from external knowledge takes a huge time and effort. This limitation severely constrains the potential of semantics in few-shot learning. In this paper, we design an automatic way called Semantic Evolution to generate high-quality semantics. The incorporation of high-quality semantics alleviates the need for complex network structures and learning algorithms used in previous works. Hence, we employ a simple two-layer network termed Semantic Alignment Network to transform semantics and visual features into robust class prototypes with rich discriminative features for few-shot classification. The experimental results show our framework outperforms all previous methods on five benchmarks, demonstrating a simple network with high-quality semantics can beat intricate multi-modal modules on few-shot classification tasks.
翻译:从有限数据中学习,即小样本学习,是一项颇具挑战性的计算机视觉任务。诸多研究利用语义信息并设计复杂的语义融合机制,以补偿有限数据中稀缺的代表性特征。然而,依赖如类别名称等朴素语义会因其简洁性引入偏差,而从外部知识获取大量语义则需耗费大量时间与精力。这一限制严重制约了语义在小样本学习中的潜力。本文设计了一种名为“语义演化”的自动方法,以生成高质量的语义信息。高质量语义的引入降低了对先前研究中复杂网络结构与学习算法的需求。因此,我们采用一个简单的两层网络——语义对齐网络,将语义与视觉特征转化为鲁棒的类别原型,其富含丰富的判别性特征,用于小样本分类。实验结果表明,我们的框架在五个基准测试中均优于所有先前方法,证明了简单的网络配合高质量语义即可在小样本分类任务中胜过复杂的多模态模块。