Recently there have been efforts to introduce new benchmark tasks for spoken language understanding (SLU), like semantic parsing. In this paper, we describe our proposed spoken semantic parsing system for the quality track (Track 1) in Spoken Language Understanding Grand Challenge which is part of ICASSP Signal Processing Grand Challenge 2023. We experiment with both end-to-end and pipeline systems for this task. Strong automatic speech recognition (ASR) models like Whisper and pretrained Language models (LM) like BART are utilized inside our SLU framework to boost performance. We also investigate the output level combination of various models to get an exact match accuracy of 80.8, which won the 1st place at the challenge.
翻译:近期,学界致力于为口语理解(SLU)引入语义解析等新型基准任务。本文描述了面向ICASSP 2023信号处理大挑战赛中口语理解大挑战赛质量赛道(赛道1)所提出的口语语义解析系统。我们针对该任务尝试了端到端与管道式两种系统方案。在SLU框架内,我们借助Whisper等强自动语音识别(ASR)模型及BART等预训练语言模型(LM)以提升性能。此外,通过探索多模型输出层级的组合策略,最终取得精确匹配准确率80.8的成绩,荣获此项挑战赛第一名。