成为VIP会员查看完整内容
VIP会员码认证
首页
主题
会员
服务
注册
·
登录
大语言模型推理
关注
2
综合
百科
VIP
热门
动态
论文
精华
Privacy from Symmetry: Orthogonally Equivariant Transformers for LLM Inference
Arxiv
0+阅读 · 6月15日
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
Arxiv
0+阅读 · 6月13日
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
Arxiv
0+阅读 · 6月14日
Communication-Efficient Verifiable Attention for LLM Inference
Arxiv
0+阅读 · 6月15日
ARMOR-MAD: Adaptive Routing for Heterogeneous Multi-Agent Debate in Large Language Model Reasoning
Arxiv
0+阅读 · 6月11日
When Does Delegation Beat Majority? A Delegation-Based Aggregator for Multi-Sample LLM Inference
Arxiv
0+阅读 · 6月11日
Personalized and Robust Proactive Robot Assistance with Uncertainty-Guided LLM Reasoning
Arxiv
0+阅读 · 6月7日
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
Arxiv
0+阅读 · 5月20日
GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference
Arxiv
0+阅读 · 5月11日
SpecSA: Bridging Speculative Decoding and Sparse Attention for Efficient LLM Inference
Arxiv
0+阅读 · 5月19日
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
Arxiv
0+阅读 · 5月4日
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
Arxiv
0+阅读 · 5月13日
Explaining Too Much? Understanding How Large Language Model Reasoning Traces Influence Performance and Metacognition
Arxiv
0+阅读 · 5月25日
The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination
Arxiv
0+阅读 · 4月17日
Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning
Arxiv
0+阅读 · 4月20日
参考链接
提示
微信扫码
咨询专知VIP会员与技术项目合作
(加微信请备注: "专知")
微信扫码咨询专知VIP会员
Top