As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- selecting the right model for each query at inference time -- has become a critical systems challenge. We present vLLM Semantic Router, a signal-driven decision routing framework for Mixture-of-Modality (MoM) model deployments. The central innovation is composable signal orchestration: the system extracts heterogeneous signal types from each request -- from sub-millisecond heuristic features (keyword patterns, language detection, context length, role-based authorization) to neural classifiers (domain, embedding similarity, factual grounding, modality) -- and composes them through configurable Boolean decision rules into deployment-specific routing policies. Different deployment scenarios -- multi-cloud enterprise, privacy-regulated, cost-optimized, latency-sensitive -- are expressed as different signal-decision configurations over the same architecture, without code changes. Matched decisions drive semantic model routing: over a dozen of selection algorithms analyze request characteristics to find the best model cost-effectively, while per-decision plugin chains enforce privacy and safety constraints (jailbreak detection, PII filtering, hallucination detection via the three-stage HaluGate pipeline). The system provides OpenAI API support for stateful multi-turn conversations, multi-endpoint and multi-provider routing across heterogeneous backends (vLLM, OpenAI, Anthropic, Azure, Bedrock, Gemini, Vertex AI), and a pluggable authorization factory supporting multiple auth providers. Deployed in production as an Envoy external processor, the architecture demonstrates that composable signal orchestration enables a single routing framework to serve diverse deployment scenarios with differentiated cost, privacy, and safety policies.


翻译:随着大语言模型在模态、能力和成本维度上的持续多样化发展,智能请求路由问题——即在推理时为每个查询选择合适模型——已成为关键的系统性挑战。我们提出vLLM语义路由器,这是一种面向混合模态模型部署的信号驱动决策路由框架。其核心创新在于可组合信号编排:系统从每个请求中提取异构信号类型——从亚毫秒级启发式特征(关键词模式、语言检测、上下文长度、基于角色的授权)到神经分类器(领域、嵌入相似性、事实基础、模态)——并通过可配置的布尔决策规则将其组合为部署特定的路由策略。不同部署场景(多云企业、隐私合规、成本优化、延迟敏感)可表达为同一架构上的不同信号-决策配置,无需修改代码。匹配的决策驱动语义模型路由:十余种选择算法分析请求特征以经济高效地寻找最佳模型,同时每个决策的插件链强制执行隐私和安全约束(越狱检测、PII过滤、通过三阶段HaluGate流水线实现的幻觉检测)。该系统提供支持有状态多轮对话的OpenAI API、跨异构后端(vLLM、OpenAI、Anthropic、Azure、Bedrock、Gemini、Vertex AI)的多端点多提供商路由,以及支持多种认证提供商的插件式授权工厂。该架构作为Envoy外部处理器在生产环境中部署,证明了可组合信号编排使单一路由框架能够以差异化的成本、隐私和安全策略服务于多种部署场景。

0
下载
关闭预览

相关内容

ACM/IEEE第23届模型驱动工程语言和系统国际会议,是模型驱动软件和系统工程的首要会议系列,由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来,模型涵盖了建模的各个方面,从语言和方法到工具和应用程序。模特的参加者来自不同的背景,包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛,参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会,并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。 官网链接:http://www.modelsconference.org/
PlanGenLLMs:大型语言模型规划能力的最新综述
专知会员服务
34+阅读 · 2025年5月18日
关于大语言模型驱动的推荐系统智能体的综述
专知会员服务
30+阅读 · 2025年2月17日
使用 OpenLLM 构建和部署大模型应用
专知会员服务
55+阅读 · 2024年1月4日
BiSeNet:双向分割网络进行实时语义分割
统计学习与视觉计算组
22+阅读 · 2018年8月23日
阿里流行音乐趋势预测-深度学习LSTM网络实现代码分享
机器学习研究会
11+阅读 · 2017年12月5日
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
VIP会员
最新内容
综述 | 终端智能体:命令行环境中的 AI Agents
专知会员服务
3+阅读 · 8月24日
综述 | 多模态智能体框架基础与前沿
专知会员服务
3+阅读 · 8月24日
《自适应无人机集群网络开发》
专知会员服务
6+阅读 · 8月24日
无人机与作战飞机最佳精确目标定位技术
专知会员服务
5+阅读 · 8月24日
博士论文 | 大动作空间中的在线与离线策略学习
论文 | 全球负责任 AI 指数 2026 方法论
专知会员服务
7+阅读 · 8月23日
超致命战场空间中的战术通信生存能力
专知会员服务
5+阅读 · 8月23日
相关基金
国家自然科学基金
6+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
1+阅读 · 2014年12月31日
国家自然科学基金
2+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
4+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员