Need for the efficient processing of neural networks has given rise to the development of hardware accelerators. The increased adoption of specialized hardware has highlighted the need for more agile design flows for hardware-software co-design and domain-specific optimizations. In this paper, we present CFU Playground: a full-stack open-source framework that enables rapid and iterative design and evaluation of machine learning (ML) accelerators for embedded ML systems. Our tool provides a completely open-source end-to-end flow for hardware-software co-design on FPGAs and future systems research. This full-stack framework gives the users access to explore experimental and bespoke architectures that are customized and co-optimized for embedded ML. Our rapid, deploy-profile-optimization feedback loop lets ML hardware and software developers achieve significant returns out of a relatively small investment in customization. Using CFU Playground's design and evaluation loop, we show substantial speedups between 55$\times$ and 75$\times$. The soft CPU coupled with the accelerator opens up a new, rich design space between the two components that we explore in an automated fashion using Vizier, an open-source black-box optimization service.
翻译:神经网络高效处理的需求催生了硬件加速器的发展。专用硬件的日益普及凸显了硬件-软件协同设计与领域特定优化中更敏捷设计流程的必要性。本文提出CFU Playground:一个全栈开源框架,能够实现嵌入式机器学习(ML)系统加速器的快速迭代设计与评估。我们的工具为FPGA上的硬件-软件协同设计及未来系统研究提供了完全开源的全流程方案。该全栈框架允许用户探索针对嵌入式ML定制与协同优化的实验性定制架构。通过部署-分析-优化的快速反馈循环,ML硬件与软件开发者能够以相对较小的定制投入获得显著回报。利用CFU Playground的设计与评估循环,我们展示了55倍至75倍的显著加速效果。软CPU与加速器的结合在两者之间开辟了新的丰富设计空间,我们通过开源黑盒优化服务Vizier以自动化方式探索该空间。