Multimodal large models have been recognized for their advantages in various performance and downstream tasks. The development of these models is crucial towards achieving general artificial intelligence in the future. In this paper, we propose a novel universal language representation learning method called UniBriVL, which is based on Bridging-Vision-and-Language (BriVL). Universal BriVL embeds audio, image, and text into a shared space, enabling the realization of various multimodal applications. Our approach addresses major challenges in robust language (both text and audio) representation learning and effectively captures the correlation between audio and image. Additionally, we demonstrate the qualitative evaluation of the generated images from UniBriVL, which serves to highlight the potential of our approach in creating images from audio. Overall, our experimental results demonstrate the efficacy of UniBriVL in downstream tasks and its ability to choose appropriate images from audio. The proposed approach has the potential for various applications such as speech recognition, music signal processing, and captioning systems.
翻译:多模态大模型因其在各类性能及下游任务中的优势而受到广泛认可,这类模型的发展对于未来实现通用人工智能至关重要。本文提出一种名为UniBriVL的新型通用语言表征学习方法,该方法基于桥接视觉与语言(BriVL)框架。通用BriVL将音频、图像和文本嵌入共享空间,从而支持多种多模态应用的实现。我们的方法攻克了鲁棒语言(包括文本和音频)表征学习中的主要挑战,并能有效捕捉音频与图像之间的相关性。此外,我们展示了UniBriVL生成图像的定性评估结果,这凸显了该方法从音频创建图像的潜力。总体而言,实验结果表明UniBriVL在下游任务中的有效性及其从音频中选择合适图像的能力。所提出的方法在语音识别、音乐信号处理和字幕系统等多种应用中具有潜在价值。