Robotic task planning in real-world environments requires reasoning over implicit constraints from language and vision. While LLMs and VLMs offer strong priors, they struggle with long-horizon structure and symbolic grounding. Existing methods that combine LLMs with symbolic planning often rely on handcrafted or narrow domains, limiting generalization. We propose UniDomain, a framework that pre-trains a PDDL domain from robot manipulation demonstrations and applies it for online robotic task planning. It extracts atomic domains from 12,393 manipulation videos to form a unified domain with 3137 operators, 2875 predicates, and 16481 causal edges. Given a target class of tasks, it retrieves relevant atomics from the unified domain and systematically fuses them into high-quality meta-domains to support compositional generalization in planning. Experiments on diverse real-world tasks show that UniDomain solves complex, unseen tasks in a zero-shot manner, achieving up to 58% higher task success and 160% improvement in plan optimality over state-of-the-art LLM and LLM-PDDL baselines.
翻译:摘要:真实环境中的机器人任务规划需对语言与视觉中的隐式约束进行推理。尽管大语言模型(LLM)与视觉语言模型(VLM)具备强大的先验知识,但在长程结构建模与符号化接地方面仍存在局限。现有结合LLM与符号规划的方法常依赖手工构建或领域狭窄的规约域,制约了泛化能力。本文提出UniDomain框架,通过机器人操作演示预训练PDDL域,并将其应用于在线机器人任务规划。该框架从12,393个操作视频中提取原子域,整合构建包含3,137个操作符、2,875个谓词及16,481条因果边的统一域。针对特定任务类,系统从统一域中检索相关原子算子,并通过系统融合生成高质量元域,以支持规划中的组合泛化。在多样化真实世界任务上的实验表明,UniDomain能以零样本方式解决复杂未见任务,任务成功率较当前最优的LLM及LLM-PDDL基线提升最高58%,规划最优性改进达160%。