This report describes our approach for the Audio-Visual Diarization (AVD) task of the Ego4D Challenge 2022. Specifically, we present multiple technical improvements over the official baselines. First, we improve the detection performance of the camera wearer's voice activity by modifying the training scheme of its model. Second, we discover that an off-the-shelf voice activity detection model can effectively remove false positives when it is applied solely to the camera wearer's voice activities. Lastly, we show that better active speaker detection leads to a better AVD outcome. Our final method obtains 65.9% DER on the test set of Ego4D, which significantly outperforms all the baselines. Our submission achieved 1st place in the Ego4D Challenge 2022.
翻译:本报告介绍了我们在2022年Ego4D挑战赛音视频说话人日志任务中采用的方法。具体而言,我们在官方基线的基础上提出了多项技术改进。首先,通过修改摄像机佩戴者语音活动检测模型的训练方案,提升了其检测性能。其次,我们发现直接对摄像机佩戴者的语音活动应用现成的语音活动检测模型,可以有效消除误检。最后,我们表明更优的主动说话人检测能够带来更佳的AVD结果。我们的最终方法在Ego4D测试集上取得了65.9%的DER,显著优于所有基线。我们的提交在2022年Ego4D挑战赛中获得了第一名。