The integration of pre-trained text-based large language models (LLM) with speech input has enabled instruction-following capabilities for diverse speech tasks. This integration requires the use of a speech encoder, a speech adapter, and an LLM, trained on diverse tasks. We propose the use of discrete speech units (DSU), rather than continuous-valued speech encoder outputs, that are converted to the LLM token embedding space using the speech adapter. We generate DSU using a self-supervised speech encoder followed by k-means clustering. The proposed model shows robust performance on speech inputs from seen/unseen domains and instruction-following capability in spoken question answering. We also explore various types of DSU extracted from different layers of the self-supervised speech encoder, as well as Mel frequency Cepstral Coefficients (MFCC). Our findings suggest that the ASR task and datasets are not crucial in instruction-tuning for spoken question answering tasks.
翻译:将基于文本的预训练大语言模型(LLM)与语音输入相结合,已使模型能够遵循指令处理多样化的语音任务。这种集成需要使用语音编码器、语音适配器和LLM,并在多样化任务上进行训练。我们提出使用离散语音单元(DSU),而非连续值的语音编码器输出,这些单元通过语音适配器转换到LLM的词元嵌入空间。我们使用自监督语音编码器配合k均值聚类来生成DSU。所提出的模型在来自已见/未见领域的语音输入上表现出鲁棒性能,并在口语问答任务中展现出指令遵循能力。我们还探索了从自监督语音编码器不同层提取的多种类型DSU,以及梅尔频率倒谱系数(MFCC)。我们的研究结果表明,自动语音识别(ASR)任务和数据集对于口语问答任务的指令微调并非至关重要。