Speech Continuation (SC) is the task of generating a coherent extension of a spoken prompt while preserving both semantic context and speaker identity. Because SC is constrained to a single audio stream, it offers a more direct setting for probing biases in speech foundation models than dialogue does. In this work we present the first systematic evaluation of bias in SC, investigating how gender and phonation type (breathy, creaky, end-creak) affect continuation behaviour. We evaluate three recent models: SpiritLM (base and expressive), VAE-GSLM, and SpeechGPT across speaker similarity, voice quality preservation, and text-based bias metrics. Results show that while both speaker similarity and coherence remain a challenge, textual evaluations reveal significant model and gender interactions: once coherence is sufficiently high (for VAE-GSLM), gender effects emerge on text-metrics such as agency and sentence polarity. In addition, continuations revert toward modal phonation more strongly for female prompts than for male ones, revealing a systematic voice-quality bias. These findings highlight SC as a controlled probe of socially relevant representational biases in speech foundation models, and suggest that it will become an increasingly informative diagnostic as continuation quality improves.
翻译:语音续写(SC)任务旨在对语音提示生成连贯的扩展内容,同时保留语义上下文和说话者身份。由于SC受限于单一音频流,相较于对话而言,它为探测语音基础模型中的偏好提供了更直接的设置。本文首次系统评估了SC中的偏好,研究了性别和发声类型(气声、嘎裂声、尾端嘎裂声)如何影响续写行为。我们评估了三种近期模型:SpiritLM(基础版与表现力版)、VAE-GSLM和SpeechGPT,从说话者相似度、音质保持和基于文本的偏好指标入手。结果表明,尽管说话者相似度和连贯性仍是挑战,但文本评估揭示了显著的模型与性别交互作用:一旦连贯性足够高(对于VAE-GSLM),性别效应会出现在文本指标(如能动性和句子极性)上。此外,相较于男性提示,女性提示更倾向于回归模态发声,揭示了系统性的音质偏好。这些发现将SC定位为探测语音基础模型中与社会相关表征偏好的受控探针,并随着续写质量提升,其将成为一个越来越具信息量的诊断工具。