The aim of steganography is to hide secret information inside ordinary media so that the existence of communication is hidden rather than encrypted. In audiovisual context, the availability of audio and video streams creates an opportunity to split a payload across these two modes thus, reducing the embedding burden on any single carrier. This paper evaluates whether such split-payload audiovisual steganography can help evade unimodal and multimodal steganalysis under synchronized and asynchronous embedding settings. We create audiovisual samples where the hidden message is divided between the audio and video tracks, and then test how well different detectors can identify them. The single mode detectors performs close to random guessing, thus showing the benefit of this hiding mechanism, while the multimodal model initially appears more effective. However, further checks show that this improvement mostly comes from the video stream, not from a true combined audio-video signal. Overall, the results suggest that splitting the payload across modalities can make detection harder, but multimodal detectors must be evaluated carefully to ensure they are learning the intended signal.
翻译:隐写术的目标是将秘密信息隐藏在普通媒体中,以隐藏通信的存在而非进行加密。在音视频背景下,音频和视频流的可用性为跨这两种模态拆分载荷创造了机会,从而减少了单个载体的嵌入负担。本文评估了这种拆分载荷音视频隐写在同步和异步嵌入设置下能否规避单模态和多模态隐写分析。我们创建了音视频样本,其中隐藏信息被分割到音频和视频轨道之间,然后测试不同检测器识别它们的能力。单模态检测器的表现接近随机猜测,从而显示了这种隐藏机制的优势,而多模态模型最初看起来更有效。然而,进一步检查表明,这种改进主要来自视频流,而非真正的音频-视频组合信号。总体而言,结果表明跨模态拆分载荷可以使检测更加困难,但必须仔细评估多模态检测器,以确保它们学习的是预期信号。