The goal of video motion magnification techniques is to magnify small motions in a video to reveal previously invisible or unseen movement. Its uses extend from bio-medical applications and deepfake detection to structural modal analysis and predictive maintenance. However, discerning small motion from noise is a complex task, especially when attempting to magnify very subtle, often sub-pixel movement. As a result, motion magnification techniques generally suffer from noisy and blurry outputs. This work presents a new state-of-the-art model based on the Swin Transformer, which offers better tolerance to noisy inputs as well as higher-quality outputs that exhibit less noise, blurriness, and artifacts than prior-art. Improvements in output image quality will enable more precise measurements for any application reliant on magnified video sequences, and may enable further development of video motion magnification techniques in new technical fields.
翻译:摘要:视频运动放大技术的目标是在视频中放大微小运动,以揭示先前不可见或无法观察到的运动。其应用涵盖生物医学、深度伪造检测、结构模态分析以及预测性维护等领域。然而,从噪声中区分微小运动是一项复杂的任务,尤其是在试图放大非常细微且常为亚像素级别的运动时。因此,现有运动放大技术通常会输出带有噪声和模糊的结果。本文提出了一种基于Swin Transformer的最新型模型,该模型对含噪声输入具有更好的容忍性,并能生成比现有技术噪声更少、模糊度更低、伪影更少的高质量输出。输出图像质量的提升将使得任何依赖放大视频序列的应用能够实现更精确的测量,并可能推动视频运动放大技术在新技术领域中的进一步发展。