Trained computer vision models are assumed to solve vision tasks by imitating human behavior learned from training labels. Most efforts in recent vision research focus on measuring the model task performance using standardized benchmarks. Limited work has been done to understand the perceptual difference between humans and machines. To fill this gap, our study first quantifies and analyzes the statistical distributions of mistakes from the two sources. We then explore human vs. machine expertise after ranking tasks by difficulty levels. Even when humans and machines have similar overall accuracies, the distribution of answers may vary. Leveraging the perceptual difference between humans and machines, we empirically demonstrate a post-hoc human-machine collaboration that outperforms humans or machines alone.
翻译:经过训练的计算机视觉模型被认为通过模仿训练标签中学习到的人类行为来解决视觉任务。近期视觉研究的主要努力集中在使用标准化基准测试来衡量模型的任务性能。但关于人类与机器之间感知差异的研究仍十分有限。为填补这一空白,本研究首先量化并分析了来自两个来源的错误统计分布,随后根据任务难度排序探索了人与机器的专长领域。即使人类与机器具有相近的总体准确率,其答案分布仍可能存在差异。利用人机感知差异,我们通过实证证明了一种后验的人机协作方法,其性能优于单独的人类或机器。