One of the fundamental skills required for an agent acting in an environment to complete tasks is the ability to understand what actions are plausible at any given point. This work explores a novel use of code representations to reason about action preconditions for sequential decision making tasks. Code representations offer the flexibility to model procedural activities and associated constraints as well as the ability to execute and verify constraint satisfaction. Leveraging code representations, we extract action preconditions from demonstration trajectories in a zero-shot manner using pre-trained code models. Given these extracted preconditions, we propose a precondition-aware action sampling strategy that ensures actions predicted by a policy are consistent with preconditions. We demonstrate that the proposed approach enhances the performance of few-shot policy learning approaches across task-oriented dialog and embodied textworld benchmarks.
翻译:智能体在环境中完成任务所需的基本技能之一,是理解在任意给定时刻哪些动作是可行的。本研究探索了一种利用代码表示推理顺序决策任务中动作前提条件的新方法。代码表示具有灵活性,能够建模过程性活动及相关约束,并具备执行与验证约束满足性的能力。借助代码表示,我们使用预训练代码模型以零样本方式从演示轨迹中提取动作前提条件。基于提取的前提条件,我们提出了一种前提感知的动作采样策略,确保策略预测的动作与前提条件一致。实验证明,所提出的方法在面向任务的对话和具身文本世界基准测试中,能够提升少样本策略学习方法的性能。