Modern search engines are built on a stack of different components, including query understanding, retrieval, multi-stage ranking, and question answering, among others. These components are often optimized and deployed independently. In this paper, we introduce a novel conceptual framework called large search model, which redefines the conventional search stack by unifying search tasks with one large language model (LLM). All tasks are formulated as autoregressive text generation problems, allowing for the customization of tasks through the use of natural language prompts. This proposed framework capitalizes on the strong language understanding and reasoning capabilities of LLMs, offering the potential to enhance search result quality while simultaneously simplifying the existing cumbersome search stack. To substantiate the feasibility of this framework, we present a series of proof-of-concept experiments and discuss the potential challenges associated with implementing this approach within real-world search systems.
翻译:现代搜索引擎构建于包含查询理解、检索、多阶段排序及问答等多个不同组件的堆栈之上,这些组件通常独立优化和部署。本文提出了一种名为“大型搜索模型”的新型概念框架,通过统一利用一个大语言模型(LLM)来重新定义传统搜索堆栈。所有任务被表述为自回归文本生成问题,允许通过使用自然语言提示对任务进行定制。该框架充分利用了LLM强大的语言理解与推理能力,不仅有望提升搜索结果质量,同时还能简化现有繁琐的搜索堆栈。为证实该框架的可行性,我们进行了一系列概念验证实验,并讨论了将这一方法应用于真实搜索系统可能面临的挑战。