Sarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, existing detection systems often rely on multimodal data, limiting their applicability in contexts where only speech is available. To address this, we propose an annotation pipeline that leverages large language models (LLMs) to generate a sarcasm dataset. Using a publicly available sarcasm-focused podcast, we employ GPT-4o and LLaMA 3 for initial sarcasm annotations, followed by human verification to resolve disagreements. We validate this approach by comparing annotation quality and detection performance on a publicly available sarcasm dataset using a collaborative gating architecture. Finally, we introduce PodSarc, a large-scale sarcastic speech dataset created through this pipeline. The detection model achieves a 73.63% F1 score, demonstrating the dataset's potential as a benchmark for sarcasm detection research.
翻译:讽刺通过语气和语境从根本上改变语义,但数据稀缺使其在语音中的检测仍面临挑战。此外,现有检测系统通常依赖多模态数据,这限制了其在仅有语音可用的场景中的应用。为解决这一问题,我们提出一种利用大型语言模型(LLMs)生成讽刺数据集的标注流程。通过公开的以讽刺为主题的播客,我们使用GPT-4o和LLaMA 3进行初始讽刺标注,随后结合人工验证解决标注分歧。我们通过在公开讽刺数据集上比较标注质量与检测性能(采用协同门控架构)来验证该方法的有效性。最后,我们引入PodSarc——通过该流程构建的大规模讽刺语音数据集。检测模型实现了73.63%的F1分数,表明该数据集可作为讽刺检测研究的基准。