This paper presents an investigation into long-tail video recognition. We demonstrate that, unlike naturally-collected video datasets and existing long-tail image benchmarks, current video benchmarks fall short on multiple long-tailed properties. Most critically, they lack few-shot classes in their tails. In response, we propose new video benchmarks that better assess long-tail recognition, by sampling subsets from two datasets: SSv2 and VideoLT. We then propose a method, Long-Tail Mixed Reconstruction, which reduces overfitting to instances from few-shot classes by reconstructing them as weighted combinations of samples from head classes. LMR then employs label mixing to learn robust decision boundaries. It achieves state-of-the-art average class accuracy on EPIC-KITCHENS and the proposed SSv2-LT and VideoLT-LT. Benchmarks and code at: tobyperrett.github.io/lmr
翻译:本文深入研究了长尾视频识别问题。我们证明,与自然采集的视频数据集及现有长尾图像基准相比,当前视频基准在多个长尾属性上存在不足。最关键的是,它们尾部缺乏少样本类别。为此,我们通过从SSv2和VideoLT两个数据集中采样子集,提出了能更好评估长尾识别性能的新型视频基准。随后,我们提出长尾混合重构方法,该方法通过将少样本类别的实例重构为头部类别样本的加权组合,减少了对这些实例的过拟合。LMR方法进一步利用标签混合技术学习鲁棒的决策边界。该方法在EPIC-KITCHENS以及新提出的SSv2-LT和VideoLT-LT基准上实现了最先进的平均类别准确率。基准与代码参见:tobyperrett.github.io/lmr