Humor generation remains challenging task for Large Language Models (LLMs), due to their subjective nature. We focus on satire, a form of humor strongly shaped by context. In this work, we present a novel pipeline for grounded satire generation that uses Retrieval-Augmented Generation (RAG) over current news to produce satirical dictionary definitions in the Finnish context. We also introduce a new task-specific evaluation framework and annotate 100 generated definitions with six human annotators, enabling analysis across multiple experimental conditions, including cultural background, source-word type, and the presence or absence of RAG. Our results show that the generated definitions are perceived as more political than humorous. Both topic-based word selection and RAG improve the political relevance of the outputs, but neither yields clear gains in humor generation. In addition, our LLM-as-a-judge evaluation of five state-of-the-art models indicates that LLMs correlate well with human judgments on political relevance, but perform poorly on humor. We release our code and annotated dataset to support further research on grounded satire generation and evaluation.
翻译:幽默生成对于大型语言模型(LLMs)而言仍是一项具有挑战性的任务,这源于其主观性本质。本研究聚焦于讽刺——一种深受语境影响的幽默形式。我们提出了一种新颖的情境化讽刺生成流水线,该流水线利用检索增强生成(RAG)技术处理实时新闻,以芬兰语境生成讽刺性词典释义。同时,我们引入了一个任务专属的新评估框架,并邀请六位人工标注者对100条生成的释义进行标注,从而实现对多种实验条件(包括文化背景、源词类型、RAG的有无)的多维度分析。结果表明,生成的释义在政治相关感知上强于幽默感。基于主题的词语选择与RAG技术均能提升输出的政治相关性,但在幽默生成方面未带来明显增益。此外,我们采用"LLM-as-a-judge"方法对五种顶尖模型进行评估,发现LLMs在政治相关性上与人类判断高度吻合,但在幽默评价上表现欠佳。我们公开了代码与标注数据集,以推动情境化讽刺生成与评估领域的深入研究。