We present a simple and generic framework for auditing a given textual conversational system, given some samples of its conversation sessions as its input. The framework computes a SWAN (Schematised Weighted Average Nugget) score based on nugget sequences extracted from the conversation sessions. Following the approaches of S-measure and U-measure, SWAN utilises nugget positions within the conversations to weight the nuggets based on a user model. We also present a schema of twenty (+1) criteria that may be worth incorporating in the SWAN framework. In our future work, we plan to devise conversation sampling methods that are suitable for the various criteria, construct seed user turns for comparing multiple systems, and validate specific instances of SWAN for the purpose of preventing negative impacts of conversational systems on users and society. This paper was written while preparing for the ICTIR 2023 keynote (to be given on July 23, 2023).
翻译:本文提出一个简单且通用的审计框架,可对给定文本对话系统进行审计,输入数据为该系统对话会话的部分样本。该框架基于从对话会话中提取的核子序列计算SWAN(模式化加权平均核子)分数。遵循S-测度与U-测度的思路,SWAN利用对话中核子的位置,依据用户模型对核子进行加权。我们还提出一个包含二十项(+1)标准的模式体系,这些标准或可纳入SWAN框架。未来工作中,我们计划设计适用于不同标准的对话采样方法,构建用于多系统对比的种子用户轮次,并验证SWAN的特定实例,以防止对话系统对用户及社会产生负面影响。本文撰写于ICTIR 2023主旨报告(拟于2023年7月23日举行)准备期间。