Rhetoric, both spoken and written, involves not only content but also style. One common stylistic tool is $\textit{parallelism}$: the juxtaposition of phrases which have the same sequence of linguistic ($\textit{e.g.}$, phonological, syntactic, semantic) features. Despite the ubiquity of parallelism, the field of natural language processing has seldom investigated it, missing a chance to better understand the nature of the structure, meaning, and intent that humans convey. To address this, we introduce the task of $\textit{rhetorical parallelism detection}$. We construct a formal definition of it; we provide one new Latin dataset and one adapted Chinese dataset for it; we establish a family of metrics to evaluate performance on it; and, lastly, we create baseline systems and novel sequence labeling schemes to capture it. On our strictest metric, we attain $F_{1}$ scores of $0.40$ and $0.43$ on our Latin and Chinese datasets, respectively.
翻译:修辞手法(无论口语还是书面语)不仅涉及内容,也涉及风格。一种常见的修辞工具是$\textit{平行结构}$(parallelism):将具有相同语言特征(如音韵、句法、语义)序列的短语并置。尽管平行结构普遍存在,自然语言处理领域却鲜有研究,错失了深入理解人类传达的结构、意义与意图本质的机会。为此,我们提出$\textit{修辞平行结构检测}$任务。我们构建了其形式化定义,提供了一个新拉丁语数据集和一个改编中文数据集,建立了评估性能的指标族,并创建了基线系统及新型序列标注方案。在最严格的评估指标下,我们在拉丁语和中文数据集上的$F_{1}$分数分别达到$0.40$和$0.43$。