The recent growth of on-device Large Language Model (LLM) inference has driven significant interest in device-edge collaborative LLM inference. As a promising architecture, Speculative Decoding (SD) is increasingly adopted where a lightweight draft model rapidly generates candidate tokens to be verified by a powerful target model. However, a fundamental challenge lies in achieving per-token resource scheduling to effectively adapt SD paradigm to resource-constrained edge environment. This paper proposes a Generative Entropy- and Lyapunov-based Adaptive Token Offloading framework, named GELATO, to maximize decoding throughput under energy constraints in a device-edge collaborative SD system. Specifically, an outer drift-plus-penalty loop makes online decisions to establish a reference drafting budget, managing long-term energy-throughput trade-off. Further, a nested entropy-driven generation mechanism executes early exiting to adapt to per-token dynamic generative uncertainty. Theoretical analysis establishes a rigorous performance bound on long-term throughput for GELATO. Extensive evaluations demonstrate that GELATO achieves a globally optimal tradeoff, outperforming state-of-the-art distributed SD architectures by 64.98% in token throughput and reducing energy consumption by 47.47% under resource-constrained environments, while preserving LLM decoding quality.


翻译:近年来,设备端大语言模型推理的快速发展极大地推动了设备-边缘协同大语言模型推理的研究。作为一种有前景的架构,推测性解码被越来越多地采用:其中轻量级草稿模型快速生成候选令牌,由强大的目标模型进行验证。然而,一个根本性挑战在于实现逐令牌资源调度,以有效适应资源受限的边缘环境中的推测性解码范式。本文提出了一种名为GELATO的基于生成熵和Lyapunov的自适应令牌卸载框架,旨在设备-边缘协同推测性解码系统中,在能量约束下最大化解码吞吐量。具体而言,外部漂移加惩罚循环进行在线决策,以建立参考草稿预算,管理长期能量与吞吐量的权衡。此外,嵌套的熵驱动生成机制执行提前退出,以适应用户逐令牌的动态生成不确定性。理论分析为GELATO建立了严格的长期吞吐量性能界限。大量评估表明,GELATO实现了全局最优权衡:在资源受限环境下,令牌吞吐量相比最先进的分布式推测性解码架构提升了64.98%,能耗降低了47.47%,同时保持了LLM的解码质量。

0
下载
关闭预览

相关内容

大模型在兵力推荐中的应用与思考
专知会员服务
34+阅读 · 2025年5月7日
大模型数学推理数据合成相关方法
专知会员服务
36+阅读 · 2025年1月19日
大语言模型算法演进综述
专知会员服务
81+阅读 · 2024年5月30日
RecInterpreter:架起大语言模型与传统推荐模型的桥梁
专知会员服务
54+阅读 · 2023年11月9日
通过集成 XNNPACK 实现推理速度飞跃
TensorFlow
26+阅读 · 2020年7月30日
因果推理学习算法资源大列表
专知
27+阅读 · 2019年3月3日
深度学习在推荐系统中的应用综述(最全)
七月在线实验室
17+阅读 · 2018年5月5日
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
VIP会员
最新内容
面向2027年及未来的海军情报改革
专知会员服务
3+阅读 · 8月5日
《无人机蜂群:释放人类-蜂群编队的潜能》
专知会员服务
6+阅读 · 8月5日
《战略战术化:一项综合性述评》
专知会员服务
5+阅读 · 8月5日
相关基金
国家自然科学基金
1+阅读 · 2015年12月31日
国家自然科学基金
0+阅读 · 2015年12月31日
国家自然科学基金
2+阅读 · 2015年12月31日
国家自然科学基金
11+阅读 · 2015年12月31日
国家自然科学基金
7+阅读 · 2015年12月31日
国家自然科学基金
12+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
国家自然科学基金
6+阅读 · 2014年12月31日
国家自然科学基金
3+阅读 · 2014年12月31日
国家自然科学基金
0+阅读 · 2014年12月31日
Top
微信扫码咨询专知VIP会员