The task of Visual Relationship Recognition (VRR) aims to identify relationships between two interacting objects in an image and is particularly challenging due to the widely-spread and highly imbalanced distribution of <subject, relation, object> triplets. To overcome the resultant performance bias in existing VRR approaches, we introduce DiffAugment -- a method which first augments the tail classes in the linguistic space by making use of WordNet and then utilizes the generative prowess of Diffusion Models to expand the visual space for minority classes. We propose a novel hardness-aware component in diffusion which is based upon the hardness of each <S,R,O> triplet and demonstrate the effectiveness of hardness-aware diffusion in generating visual embeddings for the tail classes. We also propose a novel subject and object based seeding strategy for diffusion sampling which improves the discriminative capability of the generated visual embeddings. Extensive experimentation on the GQA-LT dataset shows favorable gains in the subject/object and relation average per-class accuracy using Diffusion augmented samples.
翻译:视觉关系识别(VRR)任务旨在识别图像中两个交互对象之间的关系,由于<主体、关系、对象>三元组分布广泛且高度不平衡,该任务尤为具有挑战性。为克服现有VRR方法中由此导致的性能偏差,我们提出DiffAugment——该方法首先利用WordNet在语言空间中增强尾类,进而借助扩散模型的生成能力扩展少数类的视觉空间。我们提出一种基于每个<S,R,O>三元组困难度的新型困难感知扩散组件,并验证了困难感知扩散在生成尾类视觉嵌入方面的有效性。同时,我们提出一种基于主体和对象的扩散采样种子策略,该策略提升了生成视觉嵌入的判别能力。在GQA-LT数据集上的大量实验表明,使用扩散增强样本在主体/对象和关系的平均每类准确率上取得了显著提升。