Emerging Artificial Intelligence-enabled Internet-of-Things (AI-IoT) System-on-a-Chip (SoC) for augmented reality, personalized healthcare, and nano-robotics need to run many diverse tasks within a power envelope of a few tens of mW over a wide range of operating conditions: compute-intensive but strongly quantized Deep Neural Network (DNN) inference, as well as signal processing and control requiring high-precision floating-point. We present Marsellus, an all-digital heterogeneous SoC for AI-IoT end-nodes fabricated in GlobalFoundries 22nm FDX that combines 1) a general-purpose cluster of 16 RISC-V Digital Signal Processing (DSP) cores attuned for the execution of a diverse range of workloads exploiting 4-bit and 2-bit arithmetic extensions (XpulpNN), combined with fused MAC&LOAD operations and floating-point support; 2) a 2-8bit Reconfigurable Binary Engine (RBE) to accelerate 3x3 and 1x1 (pointwise) convolutions in DNNs; 3) a set of On-Chip Monitoring (OCM) blocks connected to an Adaptive Body Biasing (ABB) generator and a hardware control loop, enabling on-the-fly adaptation of transistor threshold voltages. Marsellus achieves up to 180 Gop/s or 3.32 Top/s/W on 2-bit precision arithmetic in software, and up to 637 Gop/s or 12.4 Top/s/W on hardware-accelerated DNN layers.
翻译:摘要:面向增强现实、个性化医疗和纳米机器人等新兴人工智能物联网(AI-IoT)系统级芯片(SoC),需在数十毫瓦功率预算内、宽泛工作条件下运行多种异构任务:包括计算密集型但强量化的深度神经网络(DNN)推理,以及需要高精度浮点运算的信号处理与控制。本文提出Marsellus——一款全数字异构AI-IoT端节点SoC,采用GlobalFoundries 22nm FDX工艺实现,集成了:1)通用计算簇,包含16个RISC-V数字信号处理(DSP)核,通过4位和2位算术扩展(XpulpNN)结合融合乘累加与加载操作及浮点支持,适配多样化工作负载;2)2至8位可重构二进制引擎(RBE),用于加速DNN中的3×3和1×1(逐点)卷积;3)片上监控(OCM)模块组,与自适应体偏置(ABB)生成器及硬件控制环路相连,实现晶体管阈值电压的在线自适应调节。Marsellus在软件实现的2位精度算术中达到180 Gop/s或3.32 Top/s/W,在硬件加速的DNN层中达到637 Gop/s或12.4 Top/s/W。