Humans possess an extraordinary ability to selectively focus on the sound source of interest amidst complex acoustic environments, commonly referred to as cocktail party scenarios. In an attempt to replicate this remarkable auditory attention capability in machines, target speaker extraction (TSE) models have been developed. These models leverage the pre-registered cues of the target speaker to extract the sound source of interest. However, the effectiveness of these models is hindered in real-world scenarios due to the potential variation or even absence of pre-registered cues. To address this limitation, this study investigates the integration of natural language to enhance the flexibility and controllability of existing TSE models. Specifically, we propose a model named LLM-TSE, wherein a large language model (LLM) to extract useful semantic cues from the user's typed text input, which can complement the pre-registered cues or work independently to control the TSE process. Our experimental results demonstrate competitive performance when only text-based cues are presented, and a new state-of-the-art is set when combined with pre-registered acoustic cues. To the best of our knowledge, this is the first work that has successfully incorporated text-based cues to guide target speaker extraction, which can be a cornerstone for cocktail party problem research.
翻译:人类在复杂声学环境(即所谓的鸡尾酒会场景)中具有非凡的能力,能够有选择性地聚焦于感兴趣的声音源。为在机器中复现这一卓越的听觉注意能力,研究者开发了目标说话人提取(TSE)模型。这些模型利用预先注册的目标说话人线索来提取感兴趣的声音源。然而,由于预先注册的线索可能存在变化甚至缺失,这些模型在现实场景中的有效性受到阻碍。为解决这一局限,本研究探索引入自然语言以增强现有TSE模型的灵活性与可控性。具体而言,我们提出名为LLM-TSE的模型,其中利用大型语言模型(LLM)从用户键入的文本输入中提取有用的语义线索,这些线索可补充预先注册的线索,或独立工作以控制TSE过程。实验结果表明,仅使用文本线索时即具备竞争力表现,而结合预先注册的声学线索时则创下新的最优性能。据我们所知,这是首项成功将文本线索用于引导目标说话人提取的工作,可为鸡尾酒会问题研究奠定基石。