The recent growth of on-device Large Language Model (LLM) inference has driven significant interest in device-edge collaborative LLM inference. As a promising architecture, Speculative Decoding (SD) is increasingly adopted where a lightweight draft model rapidly generates candidate tokens to be verified by a powerful target model. However, a fundamental challenge lies in achieving per-token resource scheduling to effectively adapt SD paradigm to resource-constrained edge environment. This paper proposes a Generative Entropy- and Lyapunov-based Adaptive Token Offloading framework, named GELATO, to maximize decoding throughput under energy constraints in a device-edge collaborative SD system. Specifically, an outer drift-plus-penalty loop makes online decisions to establish a reference drafting budget, managing long-term energy-throughput trade-off. Further, a nested entropy-driven generation mechanism executes early exiting to adapt to per-token dynamic generative uncertainty. Theoretical analysis establishes a rigorous performance bound on long-term throughput for GELATO. Extensive evaluations demonstrate that GELATO achieves a globally optimal tradeoff, outperforming state-of-the-art distributed SD architectures by 64.98% in token throughput and reducing energy consumption by 47.47% under resource-constrained environments, while preserving LLM decoding quality.
翻译:近年来,设备端大语言模型推理的快速发展极大地推动了设备-边缘协同大语言模型推理的研究。作为一种有前景的架构,推测性解码被越来越多地采用:其中轻量级草稿模型快速生成候选令牌,由强大的目标模型进行验证。然而,一个根本性挑战在于实现逐令牌资源调度,以有效适应资源受限的边缘环境中的推测性解码范式。本文提出了一种名为GELATO的基于生成熵和Lyapunov的自适应令牌卸载框架,旨在设备-边缘协同推测性解码系统中,在能量约束下最大化解码吞吐量。具体而言,外部漂移加惩罚循环进行在线决策,以建立参考草稿预算,管理长期能量与吞吐量的权衡。此外,嵌套的熵驱动生成机制执行提前退出,以适应用户逐令牌的动态生成不确定性。理论分析为GELATO建立了严格的长期吞吐量性能界限。大量评估表明,GELATO实现了全局最优权衡:在资源受限环境下,令牌吞吐量相比最先进的分布式推测性解码架构提升了64.98%,能耗降低了47.47%,同时保持了LLM的解码质量。