Non-verbal signals in speech are encoded by prosody and carry information that ranges from conversation action to attitude and emotion. Despite its importance, the principles that govern prosodic structure are not yet adequately understood. This paper offers an analytical schema and a technological proof-of-concept for the categorization of prosodic signals and their association with meaning. The schema interprets surface-representations of multi-layered prosodic events. As a first step towards implementation, we present a classification process that disentangles prosodic phenomena of three orders. It relies on fine-tuning a pre-trained speech recognition model, enabling the simultaneous multi-class/multi-label detection. It generalizes over a large variety of spontaneous data, performing on a par with, or superior to, human annotation. In addition to a standardized formalization of prosody, disentangling prosodic patterns can direct a theory of communication and speech organization. A welcome by-product is an interpretation of prosody that will enhance speech- and language-related technologies.
翻译:言语中的非言语信号通过韵律编码,承载着从会话行为到态度与情感等多重信息。尽管其重要性不言而喻,但支配韵律结构的原则尚未得到充分理解。本文提出了一套分析方案及技术概念验证,旨在对韵律信号进行分类并建立其与意义的关联。该方案对多层韵律事件的外在表征进行解读。作为实现该方案的第一步,我们提出了一种分类流程,用以解构三个阶次的韵律现象。该流程基于对预训练语音识别模型的微调,可实现多类别/多标签的同步检测,并能泛化至大量自发性言语数据,其表现与人工标注相当甚至更优。除了对韵律进行标准化形式化描述外,解构韵律模式还可引导言语组织与传播理论的发展。一个有益的副产品是对韵律的阐释将推动语言相关技术的进步。