We analyze how automatic speech recognition (ASR) errors propagate through ASR-LLM cascades in Korean spoken question answering (SQA), focusing on downstream semantic failures that conventional ASR metrics cannot fully capture. Our analysis shows that the relative downstream degradation caused by ASR errors is consistent across LLMs with different absolute performance, suggesting that cascade degradation largely tracks ASR-stage information loss. We further identify single-character Korean ASR errors as a distinct semantic-failure channel, where the gold answer becomes entirely absent from the downstream prediction despite only a minimal transcription difference. Finally, an auxiliary comparison shows that a large audio language model outperforms an ASR-LLM pipeline with a matched language backbone in noisy Korean SQA, indicating the potential of direct audio input to mitigate transcript-induced information loss.
翻译:我们分析了自动语音识别(ASR)错误如何在韩语口语问答(SQA)的ASR-LLM级联中传播,重点关注传统ASR指标无法完全捕捉的下游语义失败。我们的分析表明,由ASR错误引起的相对下游性能下降在不同绝对性能的LLM上具有一致性,这表明级联退化在很大程度上追踪了ASR阶段的信息损失。我们进一步将单字符韩语ASR错误识别为一种独特的语义失败渠道——尽管转录差异极小,但正确答案却完全从下游预测中消失。最后,一项辅助比较表明,在嘈杂的韩语SQA任务中,大型音频语言模型优于具有匹配语言骨干的ASR-LLM流程,这表明直接音频输入在缓解转录引起的信息损失方面具有潜力。