We propose TextManiA, a text-driven manifold augmentation method that semantically enriches visual feature spaces, regardless of class distribution. TextManiA augments visual data with intra-class semantic perturbation by exploiting easy-to-understand visually mimetic words, i.e., attributes. This work is built on an interesting hypothesis that general language models, e.g., BERT and GPT, encompass visual information to some extent, even without training on visual training data. Given the hypothesis, TextManiA transfers pre-trained text representation obtained from a well-established large language encoder to a target visual feature space being learned. Our extensive analysis hints that the language encoder indeed encompasses visual information at least useful to augment visual representation. Our experiments demonstrate that TextManiA is particularly powerful in scarce samples with class imbalance as well as even distribution. We also show compatibility with the label mix-based approaches in evenly distributed scarce data.
翻译:摘要:本文提出TextManiA,一种基于文本驱动的流形增强方法,能够在不受类别分布影响的情况下,从语义层面丰富视觉特征空间。TextManiA通过利用易于理解的视觉拟态词(即属性),对视觉数据进行类内语义扰动。本研究基于一个有趣的假设:通用语言模型(如BERT和GPT)即便未经过视觉训练数据的训练,也在一定程度上蕴含了视觉信息。基于该假设,TextManiA将从成熟的大型语言编码器中获得的预训练文本表示迁移至正在学习的目标视觉特征空间。我们的广泛分析表明,语言编码器确实蕴含了至少可用于增强视觉表示的视觉信息。实验证明,TextManiA在样本稀缺且类别不平衡的场景中,以及在均匀分布的数据中均表现出显著优势。我们还展示了在均匀分布的稀缺数据中,该方法与基于标签混合的方法具备兼容性。