Speech encodes multiple simultaneous attributes--linguistic content, speaker identity, dialect, gender--that conventional single-vector embeddings conflate. We present a factor-partitioned embedding framework that maps each utterance into a single vector whose subspaces correspond to distinct axes of variation. A shared acoustic encoder feeds per-axis linear projection heads, each trained via distillation from a specialist teacher or a contrastive objective over shared-label pairs. The resulting embeddings support attribute-conditioned retrieval: similarity is computed as a signed weighted sum over per-axis cosine scores, allowing retrieval that jointly considers what was said and how --or explicitly suppresses one attribute to surface another. We evaluate on cross-corpus retrieval over corpora sharing the Harvard sentence prompts, demonstrating that signed axis weighting can suppress same-speaker bias and surface semantically matched utterances across recording conditions.
翻译:语音编码了多种并行属性——语言内容、说话人身份、方言、性别——而传统的单向量嵌入将这些属性混为一谈。我们提出了一种因子分区嵌入框架,将每个语音段映射为一个单一向量,其子空间对应于不同的变异轴。共享的声学编码器为每个轴提供线性投影头,每个轴通过来自专家教师的蒸馏或基于共享标签对的对比目标进行训练。生成的嵌入支持属性条件检索:相似度计算为每个轴余弦评分的带符号加权和,从而允许联合考虑说了什么和如何说——或显式抑制一个属性以凸显另一个属性。我们在共享哈佛句子提示的语料库上进行跨语料库检索评估,展示了带符号的轴加权可以抑制同说话人偏差,并在不同录音条件下凸显语义匹配的语音段。