Tone is a crucial component of the prosody of Shanghainese, a Wu Chinese variety spoken primarily in urban Shanghai. Tone sandhi, which applies to all multi-syllabic words in Shanghainese, then, is key to natural-sounding speech. Unfortunately, recent work on Shanghainese TTS (text-to-speech) such as Apple's VoiceOver has shown poor performance with tone sandhi, especially LD (left-dominant sandhi). Here I show that word segmentation during text preprocessing can improve the quality of tone sandhi production in TTS models. Syllables within the same word are annotated with a special symbol, which serves as a proxy for prosodic information of the domain of LD. Contrary to the common practice of using prosodic annotation mainly for static pauses, this paper demonstrates that prosodic annotation can also be applied to dynamic tonal phenomena. I anticipate this project to be a starting point for bringing formal linguistic accounts of Shanghainese into computational projects. Too long have we been using the Mandarin models to approximate Shanghainese, but it is a different language with its own linguistic features, and its digitisation and revitalisation should be treated as such.
翻译:音调是上海话(一种主要在上海城区使用的吴语方言)韵律的关键组成部分。连读变调适用于上海话中所有多音节词,因此对于实现自然语音至关重要。不幸的是,近期关于上海话文本转语音(TTS)的研究(如苹果公司的VoiceOver功能)显示其在连读变调方面的表现不佳,尤其是左主导变调(LD)。本文表明,在文本预处理阶段进行分词可以提升TTS模型中连读变调生成的质量。同一单词内的音节使用特殊符号进行标注,该符号作为LD域韵律信息的代理。与通常将韵律标注主要用于静态停顿的常见做法相反,本文证明韵律标注同样可以应用于动态音调现象。我期望本项目能成为将上海话的正式语言学描述引入计算项目的起点。我们长期以来一直使用普通话模型来近似处理上海话,但上海话是一种具有自身语言学特征的独立语言,其数字化与复兴应当得到相应对待。