Transformer decoding is constrained by both attention compute and KV-cache movement. This paper presents the Ferroelectric Charge-Domain Compute Cell (FCDC), a hafnium-zirconium-oxide (HZO) memcapacitor with an access device that stores analog state nonvolatilely and performs charge-domain VMM for attention. Two deployment modes are evaluated throughout: a full-substrate mode that runs q, k, v, o projections and both attention matmuls on FCDC, and a KV-coprocessor mode that only stores KV and executes the two attention matmuls; the projection-noise budget upper-bounds the coprocessor mode. The device-to-system model is cross-checked across ngspice, CrossSim, FiPy, and NeuroSim and anchored in recent wafer-scale 10 nm HZO measurements. Across 12 pretrained LLMs (1.1-32 B dense, plus a partial-layer Mixtral-8x22B 141 B-MoE stress test at k=75% and a 128 k-context dense-Mistral replication), all-layer noise substitution adds only +2.62% WikiText-2 perplexity on Qwen3-32B and +2.90% +/- 0.33% on Mistral-7B-v0.3 (five-seed mean). End-to-end analog attention adds at most +1.68 pp on TinyLlama-1.1B and shrinks below +/-1 pp on every >=7 B model. Downstream accuracy on HellaSwag, ARC, LAMBADA, and GSM8K stays within 5% of the digital baseline for Mistral-7B (MMLU -1.6 pp). The headline energy win is nonvolatility, no refresh, and KV-cache residency. A workload-level simulator anchored on measured INT4 decode energy delivers 18-35x lower per-served-token energy on RAG and agent loops against a single-user INT4 GPU baseline; against optimized GPU serving (batched vLLM, CPU+NVMe park, power-gate) the robust advantage shrinks to 1.36-4.69x and remains >=41x on parked sessions with multi-hour residency.
翻译:Transformer解码过程同时受限于注意力计算与KV缓存搬运。本文提出铁电电荷域计算单元(FCDC),这是一种由氧化铪锆(HZO)忆电容器件与访问晶体管构成的存储单元,能够非易失性存储模拟态并实现用于注意力机制的电荷域向量矩阵乘法(VMM)。全文评估了两种部署模式:全衬底模式(在FCDC上完成q、k、v、o投影运算及两次注意力矩阵乘法)与KV协处理器模式(仅存储KV并执行两次注意力矩阵乘法),其中投影噪声预算限制了协处理器模式的应用范围。该器件-系统模型经ngspice、CrossSim、FiPy、NeuroSim交叉验证,并锚定于近期晶圆级10nm HZO测试数据。针对12个预训练大语言模型(涵盖1.1B-32B稠密模型,以及k=75%条件下的部分层Mixtral-8x22B 141B-MoE压力测试与128K上下文的稠密Mistral复制实验),全层噪声替换仅在Qwen3-32B上引入+2.62%的WikiText-2困惑度增加,在Mistral-7B-v0.3上引入+2.90%±0.33%的增加(五种子均值)。端到端模拟注意力在TinyLlama-1.1B上最多增加+1.68个百分点的困惑度,而在所有≥7B的模型中该增量降至±1个百分点以下。Mistral-7B在HellaSwag、ARC、LAMBADA、GSM8K等下游任务上的准确率与数字基线偏差保持在5%以内(MMLU偏差-1.6个百分点)。核心能效优势来源于非易失性、无需刷新操作及KV缓存驻留特性。基于实测INT4解码能量构建的工作负载级仿真器显示,在RAG与智能体循环场景中,对比单用户INT4 GPU基线,每个服务token的能量消耗降低18-35倍;对比优化后的GPU服务方案(批量vLLM、CPU+NVMe休眠、电源门控),稳健优势降至1.36-4.69倍,但在驻留数小时休眠会话场景中仍保持≥41倍的能效优势。