成为VIP会员查看完整内容
VIP会员码认证
首页
主题
会员
服务
注册
·
登录
模型服务
关注
0
综合
百科
VIP
热门
动态
论文
精华
RISE: Relay Inference and Online Scheduling for Efficient Edge-Device Collaborative Diffusion Model Services
Arxiv
0+阅读 · 6月16日
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
Arxiv
0+阅读 · 6月15日
M*: A Modular, Extensible, Serving System for Multimodal Models
Arxiv
0+阅读 · 6月13日
Tangram: Unlocking Non-Uniform KV Cache for Efficient Multi-turn LLM Serving
Arxiv
0+阅读 · 6月4日
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
Arxiv
0+阅读 · 5月30日
BIRDS: Characterizing and Understanding Biodiversity Impact of Large Language Model Serving
Arxiv
0+阅读 · 5月28日
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
Arxiv
0+阅读 · 5月20日
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
Arxiv
0+阅读 · 5月27日
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving
Arxiv
0+阅读 · 5月10日
M*: A Modular, Extensible, Serving System for Multimodal Models
Arxiv
0+阅读 · 6月10日
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving
Arxiv
0+阅读 · 4月16日
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
Arxiv
0+阅读 · 3月22日
Cache Your Prompt When It's Green: Carbon-Aware Caching for Large Language Model Serving
Arxiv
0+阅读 · 4月11日
Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving
Arxiv
0+阅读 · 4月29日
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving
Arxiv
0+阅读 · 4月17日
参考链接
提示
微信扫码
咨询专知VIP会员与技术项目合作
(加微信请备注: "专知")
微信扫码咨询专知VIP会员
Top