The field of steganography has experienced a surge of interest due to the recent advancements in AI-powered techniques, particularly in the context of multimodal setups that enable the concealment of signals within signals of a different nature. The primary objectives of all steganographic methods are to achieve perceptual transparency, robustness, and large embedding capacity - which often present conflicting goals that classical methods have struggled to reconcile. This paper extends and enhances an existing image-in-audio deep steganography method by focusing on improving its robustness. The proposed enhancements include modifications to the loss function, utilization of the Short-Time Fourier Transform (STFT), introduction of redundancy in the encoding process for error correction, and buffering of additional information in the pixel subconvolution operation. The results demonstrate that our approach outperforms the existing method in terms of robustness and perceptual transparency.
翻译:隐写术领域因近期人工智能技术的发展而备受关注,特别是在多模态设置中,能够实现将信号隐藏于不同性质的信号中。所有隐写方法的主要目标是实现感知透明性、鲁棒性和大嵌入容量——这些目标往往相互冲突,经典方法难以兼顾。本文通过关注提升鲁棒性,扩展并增强了一种现有的图像到音频深度隐写方法。所提出的改进包括损失函数的修改、短时傅里叶变换(STFT)的应用、在编码过程中引入冗余以实现纠错,以及在像素子卷积操作中缓冲附加信息。结果表明,我们的方法在鲁棒性和感知透明性方面均优于现有方法。