Controlling autonomous systems under real-world conditions often requires policies that can be evaluated with low latency and minimal energy consumption. Unfortunately, these conditions are at odds with the use of high-precision deep neural networks as controllers. In this work, we introduce Differentiable Weightless Controllers (DWCs), a symbolic-differentiable architecture that learns flexible, non-linear, yet highly efficient control policies. DWCs can be trained end-to-end via gradient-based techniques, yet compile directly into FPGA-compatible circuits with few- or even single-clock-cycle latency and nanojoule-level energy cost per action. Across five MuJoCo benchmarks, including high-dimensional Humanoid, DWCs achieve returns competitive with standard deep policies (full-precision or quantized neural networks). Furthermore, DWCs exhibit structurally sparse and interpretable connectivity patterns, enabling direct inspection of which input values influence control decisions.
翻译:在真实世界条件下控制自主系统,通常需要能够以低延迟和最小能耗进行评估的策略。遗憾的是,这些条件与使用高精度深度神经网络作为控制器相矛盾。在这项工作中,我们引入了可微分无权重控制器(DWCs),这是一种符号可微分架构,能够学习灵活、非线性且高效的策略。DWCs可通过基于梯度的技术进行端到端训练,同时可直接编译为兼容FPGA的电路,每个动作仅需几个甚至单个时钟周期的延迟和纳焦级别的能耗。在五个MuJoCo基准测试(包括高维的Humanoid)中,DWCs获得了与标准深度策略(全精度或量化神经网络)相媲美的回报。此外,DWCs展现出结构稀疏且可解释的连接模式,使得能够直接检查哪些输入值影响控制决策。