In response to innovations in machine learning (ML) models, production workloads changed radically and rapidly. TPU v4 is the fifth Google domain specific architecture (DSA) and its third supercomputer for such ML models. Optical circuit switches (OCSes) dynamically reconfigure its interconnect topology to improve scale, availability, utilization, modularity, deployment, security, power, and performance; users can pick a twisted 3D torus topology if desired. Much cheaper, lower power, and faster than Infiniband, OCSes and underlying optical components are <5% of system cost and <3% of system power. Each TPU v4 includes SparseCores, dataflow processors that accelerate models that rely on embeddings by 5x-7x yet use only 5% of die area and power. Deployed since 2020, TPU v4 outperforms TPU v3 by 2.1x and improves performance/Watt by 2.7x. The TPU v4 supercomputer is 4x larger at 4096 chips and thus ~10x faster overall, which along with OCS flexibility helps large language models. For similar sized systems, it is ~4.3x-4.5x faster than the Graphcore IPU Bow and is 1.2x-1.7x faster and uses 1.3x-1.9x less power than the Nvidia A100. TPU v4s inside the energy-optimized warehouse scale computers of Google Cloud use ~3x less energy and produce ~20x less CO2e than contemporary DSAs in a typical on-premise data center.
翻译:为应对机器学习(ML)模型的创新,生产工作负载发生了根本性且快速的变革。TPU v4是谷歌的第五代领域专用架构(DSA),也是其针对此类ML模型的第三台超级计算机。光学电路交换机(OCS)可动态重构其互连拓扑,从而提升可扩展性、可用性、利用率、模块化程度、部署效率、安全性、功耗及性能;用户可根据需要选择扭曲三维环形拓扑。相较于InfiniBand,OCS与底层光学组件成本更低、功耗更小且速度更快,其成本与功耗分别不足系统总成本和总功耗的5%与3%。每个TPU v4包含SparseCores数据流处理器,可加速依赖嵌入向量模型的运行速度达5至7倍,同时仅占用5%的芯片面积和功耗。自2020年部署以来,TPU v4性能较TPU v3提升2.1倍,能效比提升2.7倍。TPU v4超级计算机规模扩大至4096个芯片(4倍于前代),整体性能提升约10倍,配合OCS的灵活性可支撑大型语言模型。在同类规模系统中,其速度较Graphcore IPU Bow快约4.3至4.5倍,较Nvidia A100快1.2至1.7倍且功耗降低1.3至1.9倍。部署于谷歌云能效优化仓库级数据中心中的TPU v4,其能耗约为典型本地数据中心中当代DSA的1/3,二氧化碳当量排放量减少约20倍。