Semantic Retrieval (SR) has become an indispensable part of the FAQ system in the task-oriented question-answering (QA) dialogue scenario. The demands for a cross-lingual smart-customer-service system for an e-commerce platform or some particular business conditions have been increasing recently. Most previous studies exploit cross-lingual pre-trained models (PTMs) for multi-lingual knowledge retrieval directly, while some others also leverage the continual pre-training before fine-tuning PTMs on the downstream tasks. However, no matter which schema is used, the previous work ignores to inform PTMs of some features of the downstream task, i.e. train their PTMs without providing any signals related to SR. To this end, in this work, we propose an Alternative Cross-lingual PTM for SR via code-switching. We are the first to utilize the code-switching approach for cross-lingual SR. Besides, we introduce the novel code-switched continual pre-training instead of directly using the PTMs on the SR tasks. The experimental results show that our proposed approach consistently outperforms the previous SOTA methods on SR and semantic textual similarity (STS) tasks with three business corpora and four open datasets in 20+ languages.
翻译:语义检索已成为任务型问答对话场景中FAQ系统不可或缺的组成部分。近年来,电商平台或特定业务场景对跨语言智能客服系统的需求日益增长。以往研究多直接利用跨语言预训练模型进行多语言知识检索,部分工作则在下游任务微调前采用持续预训练策略。然而无论采用何种范式,既有研究均未向预训练模型传递下游任务特征——即未在预训练阶段注入与语义检索相关的信号。为此,本文提出一种基于语码转换的跨语言语义检索替代性预训练方法。我们首次将语码转换技术应用于跨语言语义检索任务,并创新性地引入语码转换持续预训练范式以替代传统直接使用预训练模型的方法。实验结果表明,在20余种语言的三个商业语料库和四个公开数据集上,本方法在语义检索与语义文本相似度任务中均持续超越现有最优方法。