Many complicated real-world tasks can be broken down into smaller, more manageable parts, and planning with prior knowledge extracted from these simplified pieces is crucial for humans to make accurate decisions. However, replicating this process remains a challenge for AI agents and naturally raises two questions: How to extract discriminative knowledge representation from priors? How to develop a rational plan to decompose complex problems? Most existing representation learning methods employing a single encoder structure are fragile and sensitive to complex and diverse dynamics. To address this issue, we introduce a multiple-encoder and individual-predictor regime to learn task-essential representations from sufficient data for simple subtasks. Multiple encoders can extract adequate task-relevant dynamics without confusion, and the shared predictor can discriminate the task characteristics. We also use the attention mechanism to generate a top-k subtask planning tree, which customizes subtask execution plans in guiding complex decisions on unseen tasks. This process enables forward-looking and globality by flexibly adjusting the depth and width of the planning tree. Empirical results on a challenging platform composed of some basic simple tasks and combinatorially rich synthetic tasks consistently outperform some competitive baselines and demonstrate the benefits of our design.
翻译:许多复杂的现实世界任务可分解为更小且更易管理的部分,利用从这些简化片段中提取的先验知识进行规划,对人类做出准确决策至关重要。然而,在人工智能代理中复现这一过程仍面临挑战,并自然引发两个问题:如何从先验中提取判别性知识表征?如何制定合理的规划以分解复杂问题?现有大多数采用单一编码器结构的表征学习方法,在面对复杂多样的动态环境时脆弱且敏感。为解决此问题,我们提出一种多编码器与独立预测器相结合的机制,从充足数据中为简单子任务学习任务关键表征。多编码器能够无混淆地提取充分的任务相关动态信息,而共享预测器则可判别任务特征。我们还利用注意力机制生成自顶k子任务规划树,该树通过定制子任务执行计划来指导未见任务上的复杂决策。通过灵活调整规划树的深度与宽度,该过程实现了前瞻性与全局性。在由若干基础简单任务与组合丰富的合成任务构成的挑战性平台上的实证结果,持续优于若干竞争性基线方法,并证明了我们设计的优势。