This paper presents a novel framework for joint speaker diarization (SD) and automatic speech recognition (ASR), named SLIDAR (sliding-window diarization-augmented recognition). SLIDAR can process arbitrary length inputs and can handle any number of speakers, effectively solving ``who spoke what, when'' concurrently. SLIDAR leverages a sliding window approach and consists of an end-to-end diarization-augmented speech transcription (E2E DAST) model which provides, locally, for each window: transcripts, diarization and speaker embeddings. The E2E DAST model is based on an encoder-decoder architecture and leverages recent techniques such as serialized output training and ``Whisper-style" prompting. The local outputs are then combined to get the final SD+ASR result by clustering the speaker embeddings to get global speaker identities. Experiments performed on monaural recordings from the AMI corpus confirm the effectiveness of the method in both close-talk and far-field speech scenarios.
翻译:本文提出了一种新颖的联合说话人日志(SD)和自动语音识别(ASR)框架,命名为SLIDAR(滑动窗口日志增强识别)。SLIDAR能够处理任意长度的输入并适应任意数量的说话人,有效同步解决“谁、在何时、说了什么”的问题。SLIDAR采用滑动窗口方法,包含一个端到端日志增强语音转录(E2E DAST)模型,该模型为每个窗口局部提供转录文本、说话人日志及说话人嵌入。E2E DAST模型基于编码器-解码器架构,并利用了序列化输出训练和“Whisper风格”提示等最新技术。通过聚类说话人嵌入以获得全局说话人身份,局部输出被整合为最终的SD+ASR结果。在AMI语料库单声道录音上的实验验证了该方法在近讲和远场语音场景中的有效性。