Generative robot policies typically begin action generation from an observation-independent standard Gaussian distribution, leaving the choice of source distribution underexplored. This work asks a simple question: where should action generation begin? We propose LeaP, a Learnable source Prior that replaces the standard Gaussian with a proprioception-conditioned diagonal Gaussian over action chunks. Parameterized by a lightweight MLP, LeaP jointly predicts the mean and state-adaptive variance of the source distribution, while keeping the downstream generator architecture and inference solver unchanged. This design provides an observation-informed yet stochastic initialization, allowing the generator to focus on precise action refinement rather than transporting samples from an uninformed noise source. On 15 RoboTwin manipulation tasks, LeaP achieves an average success rate of 81.6%, outperforming four representative baselines -- including deterministic-source methods, a no-prior counterpart, and a diffusion-bridge policy -- by 6.5 to 25.5 percentage points. The same prior consistently improves both flow-matching and diffusion-bridge generators, while using fewer parameters and converging faster. The advantage carries over to real-world deployment, where LeaP attains the best performance. These results suggest that the source distribution is an independent and reusable design axis for generative robot policies, complementary to the choice of generative dynamics.
翻译:生成式机器人策略通常从与观测无关的标准高斯分布开始动作生成,鲜有研究探索源分布的选择。本文提出一个简单问题:动作生成应从何处开始?我们提出LeaP——一种可学习先验源分布,其用本体感知条件化的对角高斯分布(作用于动作片段)替代标准高斯分布。LeaP通过轻量级MLP参数化,可联合预测源分布的均值与状态自适应方差,同时保持下游生成器架构和推理求解器不变。该设计提供了基于观测信息且具有随机性的初始化方式,使生成器能专注于精确的动作细化,而非从无信息噪声源中传输样本。在15项RoboTwin操作任务中,LeaP实现了81.6%的平均成功率,以6.5至25.5个百分点的优势超越四种代表性基线方法——包括确定性源方法、无先验对照方法及扩散桥策略。该先验能一致性地提升流匹配生成器和扩散桥生成器的性能,同时使用更少参数并加快收敛速度。该优势延续至实际部署场景,LeaP取得最优表现。这些结果表明,源分布可作为生成式机器人策略中独立且可复用的设计维度,与生成动力学机制的选择互补。