Diffusion-based planning has shown promising results in long-horizon, sparse-reward tasks by training trajectory diffusion models and conditioning the sampled trajectories using auxiliary guidance functions. However, due to their nature as generative models, diffusion models are not guaranteed to generate feasible plans, resulting in failed execution and precluding planners from being useful in safety-critical applications. In this work, we propose a novel approach to refine unreliable plans generated by diffusion models by providing refining guidance to error-prone plans. To this end, we suggest a new metric named restoration gap for evaluating the quality of individual plans generated by the diffusion model. A restoration gap is estimated by a gap predictor which produces restoration gap guidance to refine a diffusion planner. We additionally present an attribution map regularizer to prevent adversarial refining guidance that could be generated from the sub-optimal gap predictor, which enables further refinement of infeasible plans. We demonstrate the effectiveness of our approach on three different benchmarks in offline control settings that require long-horizon planning. We also illustrate that our approach presents explainability by presenting the attribution maps of the gap predictor and highlighting error-prone transitions, allowing for a deeper understanding of the generated plans.
翻译:基于扩散的规划通过在长时域、稀疏奖励任务中训练轨迹扩散模型,并利用辅助引导函数对采样轨迹进行条件约束,展现出显著成效。然而,作为生成模型,扩散模型无法保证生成可行方案,这会导致执行失败,并阻碍其在安全关键型应用中的实用性。为此,本文提出一种新方法,通过向易错方案提供优化引导来改进扩散模型生成的不可靠方案。具体而言,我们提出一种名为"修复差距"的新指标,用于评估扩散模型生成的单个方案质量。该指标由差距预测器估算,并产生修复差距引导信号,以优化扩散规划器。此外,我们引入属性映射正则化器,防止由次优差距预测器产生的对抗性优化引导,从而进一步改进不可行方案。我们在三个离线控制基准测试中验证了该方法在长时域规划场景下的有效性。同时,通过展示差距预测器的属性映射并突出易错轨迹段,我们证明了该方法具有可解释性,有助于深入理解生成方案的内在机理。