成为VIP会员查看完整内容
VIP会员码认证
首页
主题
会员
服务
注册
·
登录
NPU
关注
0
综合
百科
VIP
热门
动态
论文
精华
DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference
Arxiv
0+阅读 · 8月31日
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU
Arxiv
0+阅读 · 7月10日
Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving
Arxiv
0+阅读 · 7月17日
DAN-Scheduler: Deterministic Three-Stage Co-Optimization of Scheduling, Memory Layout, and Pipeline Overlap for General-Purpose NPUs
Arxiv
0+阅读 · 7月19日
AutoNeural: Co-Designing Vision-Language Models for NPU Inference
Arxiv
0+阅读 · 7月20日
BoltNet: An Ultra-Lightweight Convolutional Network for On-Device Plant Species Identification
Arxiv
0+阅读 · 8月12日
MINCE: Shrinking LLM Evaluation Datasets via Few-Model Monte Carlo Calibration
Arxiv
0+阅读 · 6月22日
Latency Prediction for LLM Inference on NPU Systems
Arxiv
0+阅读 · 6月17日
Ascend-RaBitQ: Heterogeneous NPU-CPU Acceleration of Billion-Scale Similarity Search with 1-bit Quantization
Arxiv
0+阅读 · 6月15日
KATANA: A Fast, Low-Power Mapping of Kalman Filters onto Edge NPUs for Real-Time Tracking
Arxiv
0+阅读 · 6月12日
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
Arxiv
0+阅读 · 6月7日
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
Arxiv
0+阅读 · 5月31日
Efficient On-Device Diffusion LLM Inference with Mobile NPU
Arxiv
0+阅读 · 6月11日
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference
Arxiv
0+阅读 · 5月22日
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
Arxiv
0+阅读 · 5月13日
参考链接
提示
微信扫码
咨询专知VIP会员与技术项目合作
(加微信请备注: "专知")
微信扫码咨询专知VIP会员
Top