The recently proposed visually grounded speech model SpeechCLIP is an innovative framework that bridges speech and text through images via CLIP without relying on text transcription. On this basis, this paper introduces two extensions to SpeechCLIP. First, we apply the Continuous Integrate-and-Fire (CIF) module to replace a fixed number of CLS tokens in the cascaded architecture. Second, we propose a new hybrid architecture that merges the cascaded and parallel architectures of SpeechCLIP into a multi-task learning framework. Our experimental evaluation is performed on the Flickr8k and SpokenCOCO datasets. The results show that in the speech keyword extraction task, the CIF-based cascaded SpeechCLIP model outperforms the previous cascaded SpeechCLIP model using a fixed number of CLS tokens. Furthermore, through our hybrid architecture, cascaded task learning boosts the performance of the parallel branch in image-speech retrieval tasks.
翻译:近期提出的视觉引导语音模型SpeechCLIP是一种创新框架,通过CLIP跨越语音与文本之间的鸿沟,无需依赖文本转录即可实现语音与图像的关联。在此基础上,本文针对SpeechCLIP提出两项扩展。首先,我们应用连续积分-触发(CIF)模块,替代级联架构中固定数量的CLS标记。其次,我们提出一种新型混合架构,将SpeechCLIP的级联与并行架构融合为多任务学习框架。实验评估基于Flickr8k和SpokenCOCO数据集展开。结果表明,在语音关键词提取任务中,基于CIF的级联SpeechCLIP模型性能优于先前使用固定数量CLS标记的级联SpeechCLIP模型。此外,通过混合架构,级联任务学习可提升并行分支在图像-语音检索任务中的表现。