Large audio-video language models can generate descriptions for both video and audio. However, they sometimes ignore audio content, producing audio descriptions solely reliant on visual information. This paper refers to this as audio hallucinations and analyzes them in large audio-video language models. We gather 1,000 sentences by inquiring about audio information and annotate them whether they contain hallucinations. If a sentence is hallucinated, we also categorize the type of hallucination. The results reveal that 332 sentences are hallucinated with distinct trends observed in nouns and verbs for each hallucination type. Based on this, we tackle a task of audio hallucination classification using pre-trained audio-text models in the zero-shot and fine-tuning settings. Our experimental results reveal that the zero-shot models achieve higher performance (52.2% in F1) than the random (40.3%) and the fine-tuning models achieve 87.9%, outperforming the zero-shot models.
翻译:大型音视频语言模型能够同时为视频和音频生成描述。然而,这些模型有时会忽略音频内容,仅凭视觉信息生成音频描述。本文将这种情况定义为音频幻觉,并对大型音视频语言模型中的此类现象展开分析。我们通过询问音频信息收集了1000个句子,并标注其中是否存在幻觉。若存在幻觉,则进一步对幻觉类型进行分类。结果显示,332个句子存在幻觉,且不同幻觉类型中名词和动词呈现显著分布规律。基于此,我们利用预训练的音频-文本模型,在零样本和微调设置下开展音频幻觉分类任务。实验结果表明,零样本模型的性能(F1值为52.2%)高于随机基准(40.3%),而微调模型以87.9%的F1值显著超越零样本模型。