Establishing evaluation schemes for spoken dialogue systems is important, but it can also be challenging. While subjective evaluations are commonly used in user experiments, objective evaluations are necessary for research comparison and reproducibility. To address this issue, we propose a framework for indirectly but objectively evaluating systems based on users' behaviours. In this paper, to this end, we investigate the relationship between user behaviours and subjective evaluation scores in social dialogue tasks: attentive listening, job interview, and first-meeting conversation. The results reveal that in dialogue tasks where user utterances are primary, such as attentive listening and job interview, indicators like the number of utterances and words play a significant role in evaluation. Observing disfluency also can indicate the effectiveness of formal tasks, such as job interview. On the other hand, in dialogue tasks with high interactivity, such as first-meeting conversation, behaviours related to turn-taking, like average switch pause length, become more important. These findings suggest that selecting appropriate user behaviours can provide valuable insights for objective evaluation in each social dialogue task.
翻译:建立口语对话系统的评估体系至关重要,但也充满挑战。虽然主观评估常用于用户实验,但客观评估对于研究比较和可重复性而言不可或缺。为解决这一问题,我们提出了一种基于用户行为进行间接但客观的系统评估框架。为此,本文研究了社交对话任务(专注倾听、求职面试和初次见面对话)中用户行为与主观评估得分之间的关系。结果表明,在用户话语占据主导地位的对话任务(如专注倾听和求职面试)中,话语数量和词汇量等指标对评估具有显著作用。观察不流畅性也能反映正式任务(如求职面试)的有效性。另一方面,在交互性强的对话任务(如初次见面对话)中,与话轮转换相关的行为(如平均切换停顿长度)变得更为重要。这些发现表明,选择合适的用户行为可为各类社交对话任务的客观评估提供有价值的见解。