In the domain of autonomous driving, the Learning from Demonstration (LfD) paradigm has exhibited notable efficacy in addressing sequential decision-making problems. However, consistently achieving safety in varying traffic contexts, especially in safety-critical scenarios, poses a significant challenge due to the long-tailed and unforeseen scenarios absent from offline datasets. In this paper, we introduce the saFety-aware strUctured Scenario representatION (FUSION), a pioneering methodology conceived to facilitate the learning of an adaptive end-to-end driving policy by leveraging structured scenario information. FUSION capitalizes on the causal relationships between decomposed reward, cost, state, and action space, constructing a framework for structured sequential reasoning under dynamic traffic environments. We conduct rigorous evaluations in two typical real-world settings of distribution shift in autonomous vehicles, demonstrating the good balance between safety cost and utility reward of FUSION compared to contemporary state-of-the-art safety-aware LfD baselines. Empirical evidence under diverse driving scenarios attests that FUSION significantly enhances the safety and generalizability of autonomous driving agents, even in the face of challenging and unseen environments. Furthermore, our ablation studies reveal noticeable improvements in the integration of causal representation into the safe offline RL problem.
翻译:在自主驾驶领域,基于演示学习(LfD)范式在解决序列决策问题中展现出显著有效性。然而,由于离线数据集中缺少长尾和不可预见场景,在不同交通环境(尤其是安全关键场景)中一致地实现安全性仍面临重大挑战。本文提出面向安全的结构化场景表示方法(FUSION),这是一种开创性方法,通过利用结构化场景信息来学习自适应端到端驾驶策略。FUSION利用分解后的奖励、成本、状态与动作空间之间的因果关系,构建了动态交通环境下结构化序列推理框架。我们在自主车辆分布偏移的两种典型现实场景中进行严格评估,结果表明相较于当代最先进的安全感知LfD基线方法,FUSION在安全成本与效用奖励之间取得了良好平衡。多样驾驶场景下的实证证据表明,即使在具有挑战性的未知环境中,FUSION也能显著增强自主驾驶智能体的安全性与泛化能力。此外,消融实验揭示了将因果表示融入安全离线强化学习问题所带来的显著改进。