Benchmarking and co-design are essential for driving optimizations and innovation around ML models, ML software, and next-generation hardware. Full workload benchmarks, e.g. MLPerf, play an essential role in enabling fair comparison across different software and hardware stacks especially once systems are fully designed and deployed. However, the pace of AI innovation demands a more agile methodology to benchmark creation and usage by simulators and emulators for future system co-design. We propose Chakra, an open graph schema for standardizing workload specification capturing key operations and dependencies, also known as Execution Trace (ET). In addition, we propose a complementary set of tools/capabilities to enable collection, generation, and adoption of Chakra ETs by a wide range of simulators, emulators, and benchmarks. For instance, we use generative AI models to learn latent statistical properties across thousands of Chakra ETs and use these models to synthesize Chakra ETs. These synthetic ETs can obfuscate key proprietary information and also target future what-if scenarios. As an example, we demonstrate an end-to-end proof-of-concept that converts PyTorch ETs to Chakra ETs and uses this to drive an open-source training system simulator (ASTRA-sim). Our end-goal is to build a vibrant industry-wide ecosystem of agile benchmarks and tools to drive future AI system co-design.
翻译:基准测试与协同设计对推动机器学习模型、软件及下一代硬件的优化与创新至关重要。完整的负载基准测试(如MLPerf)在系统完全设计与部署后,能够有效实现不同软硬件栈间的公平比较。然而,AI创新的迅猛发展要求更敏捷的基准测试方法,通过模拟器与仿真器支持未来系统的协同设计。我们提出Chakra,一种开放图模式,用于标准化工作负载规格,捕获关键操作及其依赖关系(即执行轨迹,Execution Trace, ET)。同时,我们提出一套配套工具/能力,支持Chakra ET的采集、生成与在广泛模拟器、仿真器和基准测试中的应用。例如,利用生成式AI模型学习数千条Chakra ET的潜在统计属性,并据此合成Chakra ET。这些合成ET可混淆关键专有信息,并针对未来假设场景进行模拟。作为示例,我们展示了端到端概念验证方案,将PyTorch ET转换为Chakra ET,并驱动开源训练系统模拟器(ASTRA-sim)。我们的最终目标是构建一个充满活力的行业级敏捷基准与工具生态系统,推动未来AI系统的协同设计。