Interactive search can provide a better experience by incorporating interaction feedback from the users. This can significantly improve search accuracy as it helps avoid irrelevant information and captures the users' search intents. Existing state-of-the-art (SOTA) systems use reinforcement learning (RL) models to incorporate the interactions but focus on item-level feedback, ignoring the fine-grained information found in sentence-level feedback. Yet such feedback requires extensive RL action space exploration and large amounts of annotated data. This work addresses these challenges by proposing a new deep Q-learning (DQ) approach, DQrank. DQrank adapts BERT-based models, the SOTA in natural language processing, to select crucial sentences based on users' engagement and rank the items to obtain more satisfactory responses. We also propose two mechanisms to better explore optimal actions. DQrank further utilizes the experience replay mechanism in DQ to store the feedback sentences to obtain a better initial ranking performance. We validate the effectiveness of DQrank on three search datasets. The results show that DQRank performs at least 12% better than the previous SOTA RL approaches. We also conduct detailed ablation studies. The ablation results demonstrate that each model component can efficiently extract and accumulate long-term engagement effects from the users' sentence-level feedback. This structure offers new technologies with promised performance to construct a search system with sentence-level interaction.
翻译:交互式搜索可通过融入用户交互反馈来提供更好的搜索体验。这能显著提升搜索准确率,因为它有助于过滤无关信息并捕捉用户的搜索意图。现有最优系统采用强化学习(RL)模型来整合交互行为,但主要聚焦于项目级反馈,忽略了句子级反馈中包含的细粒度信息。然而,这种反馈需要大量RL动作空间探索和大量标注数据。本研究通过提出新型深度Q学习(DQ)方法DQrank来应对这些挑战。DQrank将自然语言处理领域最先进的BERT模型进行适配,基于用户参与度选择关键句子,并对项目进行排序以获得更令人满意的响应。我们还提出了两种更优动作探索机制。DQrank进一步利用深度Q学习中的经验回放机制存储反馈句子,以获得更优的初始排序性能。我们在三个搜索数据集上验证了DQrank的有效性。结果表明DQrank比先前最先进的RL方法性能提升至少12%。我们还进行了详细的消融研究,消融结果证明每个模型组件都能从用户的句子级反馈中有效提取和累积长期参与效应。该结构为构建具有句子级交互的搜索系统提供了具有前景的新型技术。