Audio-visual learning seeks to enhance the computer's multi-modal perception leveraging the correlation between the auditory and visual modalities. Despite their many useful downstream tasks, such as video retrieval, AR/VR, and accessibility, the performance and adoption of existing audio-visual models have been impeded by the availability of high-quality datasets. Annotating audio-visual datasets is laborious, expensive, and time-consuming. To address this challenge, we designed and developed an efficient audio-visual annotation tool called Peanut. Peanut's human-AI collaborative pipeline separates the multi-modal task into two single-modal tasks, and utilizes state-of-the-art object detection and sound-tagging models to reduce the annotators' effort to process each frame and the number of manually-annotated frames needed. A within-subject user study with 20 participants found that Peanut can significantly accelerate the audio-visual data annotation process while maintaining high annotation accuracy.
翻译:视听学习旨在利用听觉与视觉模态之间的关联性来增强计算机的多模态感知能力。尽管视频检索、增强现实/虚拟现实以及无障碍技术等下游应用场景众多,但现有视听模型的性能与普及程度仍受限于高质量数据集的可用性。对视听数据集进行标注既费力、昂贵又耗时。为解决这一挑战,我们设计并开发了一种名为Peanut的高效视听标注工具。Peanut的人机协同流水线将多模态任务分解为两个单模态任务,并利用先进的物体检测与声音标注模型来减少标注者处理每帧画面的工作量以及所需手动标注的帧数。一项包含20名参与者的受试者内用户研究发现,Peanut能在保持高标注准确率的同时,显著加速视听数据标注过程。