As semiconductor power density is no longer constant with the technology process scaling down, modern CPUs are integrating capable data accelerators on chip, aiming to improve performance and efficiency for a wide range of applications and usages. One such accelerator is the Intel Data Streaming Accelerator (DSA) introduced in Intel 4th Generation Xeon Scalable CPUs (Sapphire Rapids). DSA targets data movement operations in memory that are common sources of overhead in datacenter workloads and infrastructure. In addition, it becomes much more versatile by supporting a wider range of operations on streaming data, such as CRC32 calculations, delta record creation/merging, and data integrity field (DIF) operations. This paper sets out to introduce the latest features supported by DSA, deep-dive into its versatility, and analyze its throughput benefits through a comprehensive evaluation. Along with the analysis of its characteristics, and the rich software ecosystem of DSA, we summarize several insights and guidelines for the programmer to make the most out of DSA, and use an in-depth case study of DPDK Vhost to demonstrate how these guidelines benefit a real application.
翻译:随着半导体功率密度不再随工艺制程微缩保持恒定,现代CPU开始在芯片内集成高性能数据加速器,旨在提升广泛应用程序与使用场景的性能及效率。此类加速器之一便是英特尔第四代至强可扩展CPU(Sapphire Rapids)中引入的英特尔数据流加速器(DSA)。DSA专注于内存中的数据移动操作,这些操作是数据中心工作负载与基础设施中常见的性能开销来源。此外,DSA通过支持更广泛的流式数据操作(如CRC32计算、增量记录创建/合并及数据完整性字段操作)而变得更为通用。本文旨在介绍DSA支持的最新特性,深入剖析其通用性,并通过全面评估分析其吞吐量优势。结合对其特性的分析及丰富的DSA软件生态,我们为程序员总结了若干优化洞察与使用指南,以充分发挥DSA性能,并通过DPDK Vhost的深度案例研究展示这些指南如何实际赋能应用。