Existing lip-sync deepfake detectors rely on pixel artifacts or audio-visual correspondence, and both fail under generator or language shift because the features they learn are tied to the training distribution. We take a different approach. Real lip motion is constrained by tissue mechanics and neuromuscular bandwidth; current generators impose none of these constraints, producing trajectories with elevated variance in velocity, acceleration, and jerk that real speech does not exhibit. We exploit this as a detection signal temporal lip jitter, by computing displacement, velocity, acceleration, and jerk statistics from 64 perioral landmarks over 25-frame windows and feeding them into a lightweight three-branch network. The model uses only landmark coordinates: no pixels, no audio, and no voiceprint data.
翻译:现有唇部同步深度伪造检测器依赖于像素伪影或音视频一致性,但两者在生成器或语言迁移场景下均会失效,因为它们所学习的特征与训练分布紧密绑定。我们采用不同方法。真实唇部运动受组织力学与神经肌肉带宽约束;当前生成器未施加任何此类约束,产生的运动轨迹在速度、加速度与急动度上具有超出真实语音表现的高方差。我们利用这一现象作为检测信号——即唇部时间抖动——通过计算64个口周标志点跨越25帧的时间位移、速度、加速度及急动度统计量,并将其输入轻量三分支网络。该模型仅使用标志点坐标:无需像素、音频或声纹数据。