Audio watermarking is widely used for leaking source tracing. The robustness of the watermark determines the traceability of the algorithm. With the development of digital technology, audio re-recording (AR) has become an efficient and covert means to steal secrets. AR process could drastically destroy the watermark signal while preserving the original information. This puts forward a new requirement for audio watermarking at this stage, that is, to be robust to AR distortions. Unfortunately, none of the existing algorithms can effectively resist AR attacks due to the complexity of the AR process. To address this limitation, this paper proposes DeAR, a deep-learning-based audio re-recording resistant watermarking. Inspired by DNN-based image watermarking, we pioneer a deep learning framework for audio carriers, based on which the watermark signal can be effectively embedded and extracted. Meanwhile, in order to resist the AR attack, we delicately analyze the distortions that occurred in the AR process and design the corresponding distortion layer to cooperate with the proposed watermarking framework. Extensive experiments show that the proposed algorithm can resist not only common electronic channel distortions but also AR distortions. Under the premise of high-quality embedding (SNR=25.86dB), in the case of a common re-recording distance (20cm), the algorithm can effectively achieve an average bit recovery accuracy of 98.55%.
翻译:音频水印被广泛用于泄漏源追踪,水印的鲁棒性决定了算法的可追溯性。随着数字技术的发展,音频重录(AR)已成为一种高效且隐蔽的窃密手段。重录过程在保留原始信息的同时会严重破坏水印信号,这对当前阶段的音频水印技术提出了新要求——即对重录失真具有鲁棒性。然而,由于重录过程的复杂性,现有算法均无法有效抵抗重录攻击。针对这一局限,本文提出DeAR——一种基于深度学习的抗音频重录水印方案。受基于深度神经网络的图像水印启发,我们率先为音频载体构建了深度学习框架,在此基础上可有效嵌入和提取水印信号。同时,为抵抗重录攻击,我们细致分析了重录过程中产生的失真,并设计了相应的失真层以配合所提出的水印框架。大量实验表明,该算法不仅能抵抗常见电子信道失真,还能有效抵抗重录失真。在高质量嵌入(信噪比=25.86dB)的前提下,于常见重录距离(20cm)场景中,该算法可有效达到98.55%的平均比特恢复准确率。