成为VIP会员查看完整内容
VIP会员码认证
首页
主题
会员
服务
注册
·
登录
INT8
关注
0
综合
百科
VIP
热门
动态
论文
精华
Lightweight Transformer Models for On-Device Fault Detection: A Benchmark Study on Resource-Constrained Deployment
Arxiv
0+阅读 · 6月23日
HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models
Arxiv
0+阅读 · 6月22日
SwitchBraidNet: Quantisation-Aware Lightweight Architecture for Hybrid Brain-Computer Interface
Arxiv
0+阅读 · 6月17日
SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM Serving
Arxiv
0+阅读 · 6月7日
Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs
Arxiv
0+阅读 · 6月2日
Holding the FP8 Quality Ceiling at 8-Bit Weights and Activations: INT8 and GGUF Post-Training Quantization of Ideogram 4.0 for Consumer GPUs
Arxiv
0+阅读 · 6月12日
Realizing Native INT8 Compute for Diffusion Transformers on Consumer GPUs: A Fused INT8 GEMM Kernel for Ideogram 4.0
Arxiv
0+阅读 · 6月12日
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
Arxiv
0+阅读 · 5月6日
SHARe-KAN: Post-Training Vector Quantization for Cache-Resident KAN Inference
Arxiv
0+阅读 · 4月14日
Edge AI for Automotive Vulnerable Road User Safety: Deployable Detection via Knowledge Distillation
Arxiv
0+阅读 · 4月29日
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
Arxiv
0+阅读 · 3月31日
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance
Arxiv
0+阅读 · 4月16日
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
Arxiv
0+阅读 · 4月6日
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
Arxiv
0+阅读 · 3月11日
Aeon: High-Performance Neuro-Symbolic Memory Management for Long-Horizon LLM Agents
Arxiv
0+阅读 · 2月17日
参考链接
提示
微信扫码
咨询专知VIP会员与技术项目合作
(加微信请备注: "专知")
微信扫码咨询专知VIP会员
Top