Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propose its distributed deployment to enable cooperative token generation in a multiuser edge system; its advantage is to effectively balance computational loads between resource-constrained devices and servers. The resulting architecture, termed Multi-access SPIN (Multi-SPIN), utilizes on-device small language models to generate and upload candidate token drafts, while an edge server operates the LLM to verify them in parallel batches. Given the severe heterogeneity in users' computation and communication capabilities, the draft length emerges as a critical control variable that influences node-level computation loads and multi-access latency, thereby governing the sum token goodput. Consequently, considering frequency-division multiple access, we investigate the problem of multi-access draft control, a joint optimization of draft-length control and bandwidth allocation to maximize sum token goodput. We examine two cases: (1) homogeneous draft lengths across users to facilitate server-side batching, and (2) heterogeneous draft lengths to introduce a new dimension for goodput enhancement. By developing decomposition methods, we reduce these complex optimizations into tractable sub-problems, which allow efficient draft control algorithms to be derived in closed form. Our analysis shows that the optimal bandwidth allocation compensates users with weaker computation-and-communication capabilities in the homogeneous case due to the batching synchronization requirements, whereas its heterogeneous-case counterpart rewards users with higher acceptance rates by relaxing such requirements. Experiments using Llama-2 and Qwen3.5 model pairs across diverse tasks demonstrate that Multi-SPIN improves goodput by up to 88% over heterogeneity-agnostic baselines.


翻译:投机推理(SPIN)最初被设计为一种用于加速大型语言模型的高效架构。本文提出其分布式部署方案,以实现多用户边缘系统中的协同令牌生成;其优势在于能够有效平衡资源受限设备与服务器之间的计算负载。由此产生的架构被称为多接入SPIN(Multi-SPIN),它利用设备端的小型语言模型生成并上传候选令牌草稿,同时由边缘服务器运行大型语言模型对其并行批量验证。考虑到用户计算与通信能力的严重异质性,草稿长度成为关键控制变量,它影响节点级计算负载与多接入延迟,进而控制令牌总有效吞吐率。因此,在频分多址接入场景下,我们研究了多接入草稿控制问题,即联合优化草稿长度控制与带宽分配以最大化令牌总有效吞吐率。我们考察了两种情形:(1)用户间采用同质草稿长度以促进服务器端批量处理,(2)用户间采用异质草稿长度以引入有效吞吐率提升的新维度。通过开发分解方法,我们将这些复杂优化问题简化为可处理的子问题,从而推导出闭合形式的草稿控制高效算法。分析表明:在同质情形下,由于批处理同步要求,最优带宽分配会对计算与通信能力较弱的用户进行补偿;而在异质情形下,通过放松同步要求,最优带宽分配会奖励具有更高接受率的用户。使用Llama-2与Qwen3.5模型对在多种任务上的实验表明,Multi-SPIN相比忽略异质性的基线方法可将有效吞吐率提升高达88%。

0
下载
关闭预览

相关内容

面向多模态智能的下一个Token预测:综述
专知会员服务
26+阅读 · 2024年12月30日
推荐系统融合排序的多目标寻优技术
专知会员服务
19+阅读 · 2024年8月17日
多域作战中实现边缘决策优势
专知会员服务
55+阅读 · 2024年5月31日
《边界监视多传感器融合系统中的目标跟踪》
专知会员服务
54+阅读 · 2023年6月11日
CFGAN:基于生成对抗网络的协同过滤框架
推荐算法:Match与Rank模型的交织配合
从0到1
15+阅读 · 2017年12月18日
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
《无人机脆弱性利用:网络空间力量的新域》
专知会员服务
2+阅读 · 今天4:08
美空军如何将人工智能从战场部署至后方机关
专知会员服务
11+阅读 · 7月31日
《史诗怒火行动:多域前瞻评估》49页报告
专知会员服务
7+阅读 · 7月31日
《英国防部:未来空战系统数字化战略》33页
专知会员服务
5+阅读 · 7月31日
《面向自主飞行网络的智能体人工智能架构》
专知会员服务
7+阅读 · 7月31日
“史诗怒火”行动:现代多域作战的重要节点
专知会员服务
8+阅读 · 7月30日
《下一代无线网络中的多无人机通信资源管理》
相关VIP内容
相关基金
国家自然科学基金
1+阅读 · 2017年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员