Query understanding in large-scale industrial search systems is typically implemented as a cascade of disparate, task-specific components. While individually optimizable, this fragmented architecture incurs high maintenance overhead and results in inconsistent behaviors, particularly for long-tail queries. In this work, we propose and deploy a unified structured query understanding system that consolidates these heterogeneous functions into a single Small Language Model (SLM) that performs schema-constrained generation. To address the data bottlenecks inherent in unified modeling, we introduce Query Illuminator, a dual-purpose framework serving as: (i) a teacher model for high-quality auto-annotation and distillation, and (ii) a surrogate judge for scalable evaluation where human labels are scarce. We validate this approach through extensive offline and online tests within LinkedIn's Job Search system. Furthermore, we demonstrate the framework's horizontal extensibility through a cross-domain case study on People Search. The results show improved user engagement and reduced operational costs, achieved while satisfying strict low-latency serving constraints on limited GPU resources.
翻译:在大规模工业级搜索系统中,查询理解通常以一组相互独立、任务特定的串行化组件方式实现。尽管每个组件可单独优化,但这种碎片化架构会导致高昂的维护开销及不一致的行为表现,尤其对于长尾查询而言。本文提出并部署了一套统一的结构化查询理解系统,将上述异构功能整合至单个执行模式约束生成的小语言模型(SLM)中。针对统一建模固有的数据瓶颈问题,我们引入Query Illuminator——一个兼具双重功能的框架:其一,作为教师模型实现高质量自动标注与知识蒸馏;其二,在人工标注稀缺的场景下充当替代评判器进行可扩展的评估。通过在LinkedIn求职搜索系统中开展的广泛离线与在线测试,我们验证了该方法的有效性。进一步地,我们通过跨领域案例研究(人物搜索)展示了该框架的水平扩展能力。结果表明,该方法在受限GPU资源上满足严格低延迟服务约束的前提下,提升了用户参与度并降低了运营成本。