3D Convolutional Neural Networks are gaining increasing attention from researchers and practitioners and have found applications in many domains, such as surveillance systems, autonomous vehicles, human monitoring systems, and video retrieval. However, their widespread adoption is hindered by their high computational and memory requirements, especially when resource-constrained systems are targeted. This paper addresses the problem of mapping X3D, a state-of-the-art model in Human Action Recognition that achieves accuracy of 95.5\% in the UCF101 benchmark, onto any FPGA device. The proposed toolflow generates an optimised stream-based hardware system, taking into account the available resources and off-chip memory characteristics of the FPGA device. The generated designs push further the current performance-accuracy pareto front, and enable for the first time the targeting of such complex model architectures for the Human Action Recognition task.
翻译:三维卷积神经网络正越来越受到研究者和从业者的关注,并在监控系统、自动驾驶车辆、人体监测系统以及视频检索等多个领域得到应用。然而,其广泛采用受到高计算和内存需求的阻碍,尤其是面向资源受限系统时。本文解决了将X3D(一种在UCF101基准测试中达到95.5%准确率的人体动作识别最新模型)映射至任意FPGA器件上的问题。所提出的工具流可根据FPGA器件的可用资源和片外存储器特性,生成优化的基于流的硬件系统。生成的硬件设计进一步推动了当前性能-精度帕累托前沿,并首次实现了针对人体动作识别任务中此类复杂模型架构的部署。