We propose VADER, a spatio-temporal matching, alignment, and change summarization method to help fight misinformation spread via manipulated videos. VADER matches and coarsely aligns partial video fragments to candidate videos using a robust visual descriptor and scalable search over adaptively chunked video content. A transformer-based alignment module then refines the temporal localization of the query fragment within the matched video. A space-time comparator module identifies regions of manipulation between aligned content, invariant to any changes due to any residual temporal misalignments or artifacts arising from non-editorial changes of the content. Robustly matching video to a trusted source enables conclusions to be drawn on video provenance, enabling informed trust decisions on content encountered.
翻译:我们提出VADER,一种时空匹配、对齐与变更摘要方法,旨在对抗通过篡改视频传播的误导性信息。VADER利用鲁棒视觉描述符与基于自适应分块视频内容的可扩展搜索,实现部分视频片段与候选视频的匹配及粗粒度对齐。随后,基于Transformer的对齐模块精细化查询片段在匹配视频中的时序定位。时空比较器模块识别对齐内容间的篡改区域,该过程不受残余时序错位或非编辑性内容变更引起的伪影影响。通过将视频鲁棒匹配至可信来源,可对视频溯源作出推断,从而为内容可信度决策提供依据。