Preference-based reinforcement learning (PbRL) is the dominant framework for aligning AI systems to human preferences. However, evaluation protocols for such data were designed for text and have not been validated for speech. We present the first ICC-based, controlled cross-modal study of human and synthetic preference annotations, comparing text and audio evaluations of identical semantic content across 100 prompts. We show that achieving $\textit{good}$ agreement within either modality (ICC(2,$k$) $\approx$ .80) requires $\sim$9 raters. At the same time, modalities show marked differences in how people report preferences: audio raters exhibit narrower decision thresholds, reduced length bias, and more user-oriented evaluation criteria, with near-chance cross-modality agreement. We demonstrate that synthetic ratings can be used to effectively predict inter-rater agreement, thus serving as an early signal for stimulus selection and proxy for human annotations. Together, these findings argue that evaluation protocols for audio preference data require modality-specific design rather than direct adaptation from text.
翻译:基于偏好的强化学习(PbRL)是使AI系统与人类偏好对齐的主流框架。然而,此类数据的评估协议是为文本设计的,尚未在语音领域得到验证。我们首次提出了基于ICC的、受控的跨模态人工与合成偏好标注对比研究,针对100个提示中语义内容完全相同的文本与音频评估进行比较。研究表明,在单一模态内达成$\textit{良好}$一致性(ICC(2,$k$) $\approx$ .80)需要约9名评分者。同时,模态间在人们报告偏好的方式上表现出显著差异:音频评分者表现出更窄的决策阈值、更弱的长度偏差以及更以用户为导向的评估标准,跨模态一致性近乎随机。我们证明合成评分可有效预测评分者间一致性,从而作为刺激选择的早期信号及人工标注的代理。综上,这些发现表明音频偏好数据的评估协议需要特定于模态的设计,而非简单地从文本领域直接迁移。