We present a temporally layered architecture (TLA) for temporally adaptive control with minimal energy expenditure. The TLA layers a fast and a slow policy together to achieve temporal abstraction that allows each layer to focus on a different time scale. Our design draws on the energy-saving mechanism of the human brain, which executes actions at different timescales depending on the environment's demands. We demonstrate that beyond energy saving, TLA provides many additional advantages, including persistent exploration, fewer required decisions, reduced jerk, and increased action repetition. We evaluate our method on a suite of continuous control tasks and demonstrate the significant advantages of TLA over existing methods when measured over multiple important metrics. We also introduce a multi-objective score to qualitatively assess continuous control policies and demonstrate a significantly better score for TLA. Our training algorithm uses minimal communication between the slow and fast layers to train both policies simultaneously, making it viable for future applications in distributed control.
翻译:我们提出了一种时间分层架构(TLA),用于以最小能耗实现时间自适应控制。TLA 通过分层整合快速与慢速策略,实现时间抽象,使每层专注于不同的时间尺度。我们的设计借鉴了人脑的节能机制,即根据环境需求在不同时间尺度上执行动作。我们证明,除节能外,TLA 还提供了诸多额外优势,包括持续探索、减少决策次数、降低加加速度以及增加动作重复率。我们在连续控制任务集上评估了该方法,并展示了 TLA 在多个关键指标上相较于现有方法的显著优势。我们还引入了一种多目标评分来定性评估连续控制策略,并证明 TLA 可获得明显更优的评分。我们的训练算法在慢速与快速层之间仅需最小通信即可同时训练两种策略,使其适用于未来的分布式控制应用。