Establishing evaluation schemes for spoken dialogue systems is important, but it can also be challenging. While subjective evaluations are commonly used in user experiments, objective evaluations are necessary for research comparison and reproducibility. To address this issue, we propose a framework for indirectly but objectively evaluating systems based on users' behaviors. In this paper, to this end, we investigate the relationship between user behaviors and subjective evaluation scores in social dialogue tasks: attentive listening, job interview, and first-meeting conversation. The results reveal that in dialogue tasks where user utterances are primary, such as attentive listening and job interview, indicators like the number of utterances and words play a significant role in evaluation. Observing disfluency also can indicate the effectiveness of formal tasks, such as job interview. On the other hand, in dialogue tasks with high interactivity, such as first-meeting conversation, behaviors related to turn-taking, like average switch pause length, become more important. These findings suggest that selecting appropriate user behaviors can provide valuable insights for objective evaluation in each social dialogue task.
翻译:建立口语对话系统的评估方案至关重要,但同时也颇具挑战性。虽然主观评估在用户实验中广泛应用,但为了研究对比和可复现性,客观评估不可或缺。为解决这一问题,我们提出了一种基于用户行为间接但客观地评估系统的框架。为此,本文研究了社交对话任务(包括专注聆听、求职面试和初次见面对话)中用户行为与主观评估得分之间的关系。结果表明,在用户话语占主导的对话任务(如专注聆听和求职面试)中,话语数和词数等指标对评估有显著影响。观察不流畅性也能反映正式任务(如求职面试)的有效性。另一方面,在高交互性的对话任务(如初次见面对话)中,与话轮转换相关的行为(如平均切换停顿长度)变得更为重要。这些发现表明,选择合适的用户行为可为各类社交对话任务的客观评估提供有价值的见解。