Audio watermarking is widely used for leaking source tracing. The robustness of the watermark determines the traceability of the algorithm. With the development of digital technology, audio re-recording (AR) has become an efficient and covert means to steal secrets. AR process could drastically destroy the watermark signal while preserving the original information. This puts forward a new requirement for audio watermarking at this stage, that is, to be robust to AR distortions. Unfortunately, none of the existing algorithms can effectively resist AR attacks due to the complexity of the AR process. To address this limitation, this paper proposes DeAR, a deep-learning-based audio re-recording resistant watermarking. Inspired by DNN-based image watermarking, we pioneer a deep learning framework for audio carriers, based on which the watermark signal can be effectively embedded and extracted. Meanwhile, in order to resist the AR attack, we delicately analyze the distortions that occurred in the AR process and design the corresponding distortion layer to cooperate with the proposed watermarking framework. Extensive experiments show that the proposed algorithm can resist not only common electronic channel distortions but also AR distortions. Under the premise of high-quality embedding (SNR=25.86dB), in the case of a common re-recording distance (20cm), the algorithm can effectively achieve an average bit recovery accuracy of 98.55%.
翻译:音频水印广泛用于泄密溯源。水印的鲁棒性决定了算法的可追溯能力。随着数字技术的发展,音频重录制已成为窃取机密的高效隐蔽手段。重录制过程在保留原始信息的同时会严重破坏水印信号,这对现阶段音频水印提出了新的要求——具备抵抗重录制失真的鲁棒性。然而受限于重录制过程的复杂性,现有算法均无法有效抵御重录制攻击。针对这一局限,本文提出DeAR——一种基于深度学习的抗音频重录制水印方案。受深度神经网络图像水印的启发,我们率先构建了面向音频载体的深度学习框架,基于该框架可实现水印信号的有效嵌入与提取。同时,为抵御重录制攻击,我们精细分析了重录制过程中产生的失真,并设计相应失真层以配合所提水印框架。大量实验表明,所提算法不仅能抵抗常见电子信道失真,更能有效抵抗重录制失真。在高品质嵌入(信噪比=25.86dB)前提下,在常见录制距离(20cm)条件下,该算法可实现平均98.55%的比特恢复准确率。