Authorship verification is the problem of determining if two distinct writing samples share the same author and is typically concerned with the attribution of written text. In this paper, we explore the attribution of transcribed speech, which poses novel challenges. The main challenge is that many stylistic features, such as punctuation and capitalization, are not available or reliable. Therefore, we expect a priori that transcribed speech is a more challenging domain for attribution. On the other hand, other stylistic features, such as speech disfluencies, may enable more successful attribution but, being specific to speech, require special purpose models. To better understand the challenges of this setting, we contribute the first systematic study of speaker attribution based solely on transcribed speech. Specifically, we propose a new benchmark for speaker attribution focused on conversational speech transcripts. To control for spurious associations of speakers with topic, we employ both conversation prompts and speakers' participating in the same conversation to construct challenging verification trials of varying difficulties. We establish the state of the art on this new benchmark by comparing a suite of neural and non-neural baselines, finding that although written text attribution models achieve surprisingly good performance in certain settings, they struggle in the hardest settings we consider.
翻译:作者身份验证是判断两个不同写作样本是否属于同一作者的问题,通常涉及书面文本的归属。本文探讨了转录语音的归属问题,这带来了新的挑战。主要挑战在于许多文体特征(如标点符号和大写字母)不可用或不可靠。因此,我们先验地预期转录语音对归属而言是一个更具挑战性的领域。另一方面,其他文体特征(如言语不流畅)可能有助于更成功的归属,但由于其语音特异性,需要专门的模型。为了更好地理解这一情境下的挑战,我们首次基于转录语音对说话者归属进行了系统研究。具体而言,我们提出了一个专注于对话语音转录的新基准。为了控制说话者与主题的虚假关联,我们利用对话提示以及说话者参与同一对话,构建了难度各异的挑战性验证试验。通过比较一系列神经网络和非神经网络基线方法,我们在这个新基准上建立了当前最优性能,并发现在某些设置中,书面文本归属模型尽管取得了惊人的良好表现,但在我们考虑的最困难设置中仍存在困难。