We investigate the integration of Large Language Models (LLMs) into query encoders to improve dense retrieval without increasing latency and cost, by circumventing the dependency on LLMs at inference time. SoftQE incorporates knowledge from LLMs by mapping embeddings of input queries to those of the LLM-expanded queries. While improvements over various strong baselines on in-domain MS-MARCO metrics are marginal, SoftQE improves performance by 2.83 absolute percentage points on average on five out-of-domain BEIR tasks.
翻译:我们研究将大语言模型(LLMs)整合到查询编码器中,通过规避推理阶段对LLMs的依赖来提升稠密检索性能,且不增加延迟和成本。SoftQE通过将输入查询的嵌入映射至LLM扩展查询的嵌入,从而融入LLMs的知识。尽管在领域内MS-MARCO指标上相较多种强基线模型的提升幅度有限,但SoftQE在五项域外BEIR任务上平均提升了2.83个绝对百分点。