Metaphor is a prominent linguistic device in human language and literature, as they add color, imagery, and emphasis to enhance effective communication. This paper introduces a large-scale high quality annotated Chinese Metaphor Corpus, which comprises around 28K sentences drawn from a diverse range of Chinese literary sources, such as poems, prose, song lyrics, etc. To ensure the accuracy and consistency of our annotations, we introduce a comprehensive set of guidelines. These guidelines address the facets of metaphor annotation, including identifying tenors, vehicles, and grounds to handling the complexities of similes, personifications, juxtapositions, and hyperboles. Breaking tradition, our approach to metaphor generation emphasizes grounds and their distinct features rather than the conventional combination of tenors and vehicles. By integrating "ground" as a CoT (Chain of Thoughts) input, we are able to generate metaphors that resonate more with real-world intuition. We test generative models such as Belle, Baichuan, and Chinese-alpaca-33B using our annotated corpus. These models are able to generate creative and fluent metaphor sentences more frequently induced by selected samples from our dataset, demonstrating the value of our corpus for Chinese metaphor research. The code is available in the https://anonymous.4open.science/r/Chinese_Metaphor_Explanation-63F2.
翻译:隐喻是人类语言与文学中重要的修辞手段,通过增添色彩、意象和强调来增强有效沟通。本文提出一个大规模高质量标注的中文隐喻语料库,包含约2.8万个句子,这些句子选自诗歌、散文、歌词等多元中文文学资源。为确保标注的准确性和一致性,我们制定了一套全面的标注指南。该指南针对隐喻标注的多个层面进行了规范,包括识别本体、喻体和喻底,并处理明喻、拟人、并列和夸张等复杂修辞现象。与传统方法不同,我们的隐喻生成方法强调喻底及其独特特征,而非单纯依赖本体与喻体的常规组合。通过将"喻底"作为思维链输入,我们能够生成更符合现实直觉的隐喻。我们使用标注语料库测试了Belle、Baichuan和Chinese-alpaca-33B等生成模型,这些模型通过数据集中的精选样本引导,能更频繁地生成富有创意且流畅的隐喻句子,证明了该语料库对中文隐喻研究的价值。代码已开源在https://anonymous.4open.science/r/Chinese_Metaphor_Explanation-63F2。