In this paper, we focus on solving one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Recent works have complex designs such as compressing the output temporally for the speech encoder, tackling modal alignment for the projector, and utilizing parameter-efficient fine-tuning for the LLM. We found that delicate designs are not necessary, while an embarrassingly simple composition of off-the-shelf speech encoder, LLM, and the only trainable linear projector is competent for the ASR task. To be more specific, we benchmark and explore various combinations of LLMs and speech encoders, leading to the optimal LLM-based ASR system, which we call SLAM-ASR. The proposed SLAM-ASR provides a clean setup and little task-specific design, where only the linear projector is trained. To the best of our knowledge, SLAM-ASR achieves the best performance on the Librispeech benchmark among LLM-based ASR models and even outperforms the latest LLM-based audio-universal model trained on massive pair data. Finally, we explore the capability emergence of LLM-based ASR in the process of modal alignment. We hope that our study can facilitate the research on extending LLM with cross-modality capacity and shed light on the LLM-based ASR community.
翻译:本文聚焦于解决语音处理领域最重要的任务之一——自动语音识别(ASR),采用语音基础编码器与大语言模型(LLM)相结合的方式。现有研究设计了复杂方案,例如对语音编码器的输出进行时间维度压缩、解决投影器的模态对齐问题,以及利用参数高效微调技术调整LLM。我们发现精细设计并非必要,仅需将现成语音编码器、LLM与唯一可训练的线性投影器进行极其简单的组合,即可胜任ASR任务。具体而言,我们基准测试并探索了多种LLM与语音编码器的组合方案,最终构建出最优的基于LLM的ASR系统,命名为SLAM-ASR。该方案采用简洁框架,几乎没有任务特定设计,仅训练线性投影器。据我们所知,SLAM-ASR在Librispeech基准测试中取得了基于LLM的ASR模型最佳性能,甚至超越了最新基于LLM、在大规模配对数据上训练的多模态音频通用模型。最后,我们探索了模态对齐过程中基于LLM的ASR涌现能力。希望本研究能推动跨模态能力扩展的LLM研究,并为基于LLM的ASR领域提供启示。