Existing conversational search studies mainly focused on asking better clarifying questions and/or improving search result quality. These works aim at retrieving better responses according to the search context, and their performances are evaluated on either single-turn tasks or multi-turn tasks under naive conversation policy settings. This leaves some questions about their applicability in real-world multi-turn conversations where realistically, each and every action needs to be made by the system itself, and search session efficiency is often an important concern of conversational search systems. While some recent works have identified the need for improving search efficiency in conversational search, they mostly require extensive data annotations and use hand-crafted rewards or heuristics to train systems that can achieve reasonable performance in a restricted number of turns, which has limited generalizability in practice. In this paper, we propose a reward-free conversation policy imitation learning framework, which can train a conversation policy without annotated conversation data or manually designed rewards. The trained conversation policy can be used to guide the conversational retrieval models to balance conversational search quality and efficiency. To evaluate the proposed conversational search system, we propose a new multi-turn-multi-response conversational evaluation metric named Expected Conversational Reciprocal Rank (ECRR). ECRR is designed to evaluate entire multi-turn conversational search sessions towards comprehensively evaluating both search result quality and search efficiency.
翻译:现有对话搜索研究主要关注提出更好的澄清问题和/或提升搜索结果质量。这些工作旨在根据搜索上下文检索更优的回应,其性能评估通常基于单轮任务或在简单对话策略设置下的多轮任务。这引发了关于其在真实多轮对话中适用性的疑问——实际场景中,每个动作都必须由系统自主完成,而搜索会话效率往往是对话搜索系统的重要考量。尽管近期部分研究已认识到提升对话搜索效率的必要性,但大多需要大量数据标注,并依赖人工设计的奖励或启发式方法训练系统,以在有限轮次内实现合理性能,这限制了在实际中的泛化能力。本文提出一种无奖励的对话策略模仿学习框架,该框架无需标注对话数据或人工设计奖励即可训练对话策略。训练后的对话策略可用于引导对话式检索模型,平衡搜索质量与效率。为评估所提对话搜索系统,我们提出一种新的多轮多响应对话评价指标——期望对话倒数排名(ECRR)。ECRR旨在评估完整的多轮对话搜索会话,以综合衡量搜索结果质量与搜索效率。