Textual humor is enormously diverse and computational studies need to account for this range, including intentionally bad humor. In this paper, we curate and analyze a novel corpus of sentences from the Bulwer-Lytton Fiction Contest to better understand "bad" humor in English. Standard humor detection models perform poorly on our corpus, and an analysis of literary devices finds that these sentences combine features common in existing humor datasets (e.g., puns, irony) with metaphor, metafiction and simile. LLMs prompted to synthesize contest-style sentences imitate the form but exaggerate the effect by over-using certain literary devices, and including far more novel adjective-noun bigrams than human writers. Data, code and analysis are available at https://github.com/venkatasg/bulwer-lytton
翻译:文本幽默具有极大的多样性,计算研究需要涵盖这一范围,包括刻意为之的低俗幽默。本文整理并分析了一个来自布尔沃-利顿小说竞赛的新颖句子语料库,以更好地理解英语中的“劣质”幽默。标准幽默检测模型在该语料库上表现不佳,而对文学手法的分析发现,这些句子将现有幽默数据集常见特征(如双关、反讽)与隐喻、元小说和明喻相结合。被提示生成类似竞赛风格句子的大型语言模型虽能模仿其形式,但会过度使用特定文学手法而夸大效果,且生成的形容词-名词双词短语数量远超人类作者。数据、代码及分析见 https://github.com/venkatasg/bulwer-lytton