Hearing is arguably an essential ability of artificial intelligence (AI) agents in the physical world, which refers to the perception and understanding of general auditory information consisting of at least three types of sounds: speech, audio events, and music. In this paper, we propose SALMONN, a speech audio language music open neural network, built by integrating a pre-trained text-based large language model (LLM) with speech and audio encoders into a single multimodal model. SALMONN enables the LLM to directly process and understand general audio inputs and achieve competitive performances on a number of speech and audio tasks used in training, such as automatic speech recognition and translation, auditory-information-based question answering, emotion recognition, speaker verification, and music and audio captioning \textit{etc.} SALMONN also has a diverse set of emergent abilities unseen in the training, which includes but is not limited to speech translation to untrained languages, speech-based slot filling, spoken-query-based question answering, audio-based storytelling, and speech audio co-reasoning \textit{etc}. The presence of the cross-modal emergent abilities is studied, and a novel few-shot activation tuning approach is proposed to activate such abilities of SALMONN. To our knowledge, SALMONN is the first model of its type and can be regarded as a step towards AI with generic hearing abilities. An interactive demo of SALMONN is available at \texttt{\url{https://github.com/bytedance/SALMONN}}, and the training code and model checkpoints will be released upon acceptance.
翻译:听觉可以说是人工智能体在物理世界中一项至关重要的能力,它涉及对至少包含三种声音类型(语音、音频事件和音乐)的通用听觉信息的感知与理解。本文提出SALMONN(语音-音频-语言-音乐开放神经网络),该模型通过将预训练的基于文本的大语言模型与语音和音频编码器集成到一个多模态模型中构建而成。SALMONN使大语言模型能够直接处理和理解通用音频输入,并在训练所涉及的多个语音和音频任务上取得有竞争力的表现,例如自动语音识别与翻译、基于听觉信息的问题回答、情感识别、说话人验证、音乐与音频描述等。SALMONN还展现出训练中未曾出现的多种涌现能力,包括但不限于:对未训练语言的语音翻译、基于语音的槽位填充、基于口语查询的问题回答、基于音频的故事生成,以及语音-音频协同推理等。本文研究了跨模态涌现能力的产生机制,并提出了一种新颖的少样本激活微调方法,以激发SALMONN的此类能力。据我们所知,SALMONN是首个此类模型,可视为迈向具备通用听觉能力的人工智能体的一步。SALMONN的交互式演示可访问\texttt{\url{https://github.com/bytedance/SALMONN}},训练代码和模型检查点将在论文被接收后发布。