Skill libraries in deployed robotic systems are continually updated through fine-tuning, fresh demonstrations, or domain adaptation, yet existing typed-composition methods (BLADE, SymSkill, Generative Skill Chaining) treat the library as frozen at test time and do not analyze how composition outcomes change when a skill is replaced. We introduce a paired-sampling cross-version swap protocol on robosuite manipulation tasks to characterize this dimension of compositional skill learning. On a dual-arm peg-in-hole task we discover a dominant-skill effect: one ECM achieves 86.7% atomic success rate while every other ECM is at or below 26.7%, and whether this dominant ECM enters a composition shifts the success rate by up to +50pp. We characterize the boundary on a simpler pick task where all atomic policies saturate at 100% and the effect is undefined. Across three tasks we further find that off-policy behavioral distance metrics fail to identify the dominant ECM, ruling out the natural cheap predictor. We propose an atomic-quality probe and a Hybrid Selector combining per-skill probes (zero per-decision cost) with selective composition revalidation (full cost), and characterize its Pareto frontier on 144 skill-update decisions. On T6 the atomic-only probe sits 23pp below full revalidation (64.6% vs 87.5% oracle match) at zero per-decision cost; a Hybrid Selector with m=10 closes most of that gap to ~12pp at 46% of full-revalidation cost. On the cross-task average over 144 events, atomic-only is within 3pp of full revalidation under a mixed-oracle caveat. The atomic-quality probe is, to our knowledge, the first principled, deployment-ready primitive for skill-update governance in compositional robot policies.
翻译:在已部署机器人系统的技能库中,技能通过微调、新示教或领域自适应不断更新,但现有类型化组合方法(如BLADE、SymSkill、Generative Skill Chaining)将技能库视为测试时冻结的静态集合,未分析替换某个技能后组合结果的变化。针对这一维度,我们在robosuite操作任务中引入配对采样跨版本交换协议,以刻画组合式技能学习的特性。在双臂插销孔任务中,我们发现了主导技能效应:一个深度可学习控制器(ECM)实现了86.7%的原子成功率,而其他所有ECM的成功率均不高于26.7%,该主导ECM是否参与组合可使成功率波动高达+50个百分点。在更为简单的抓取任务中(所有原子策略均饱和于100%),该效应的边界条件无法定义。进一步研究发现,在三个任务中,离策略行为距离指标均无法识别主导ECM,排除了天然廉价预测器的可能性。我们提出原子质量探针与混合选择器,该选择器结合了零决策成本的逐技能探针与全成本的组合重新验证机制,并在144项技能更新决策中刻画了其帕累托前沿。实验显示:在T6任务中,原子探针以零决策成本达到64.6%的匹配率(与全重新验证的87.5%相比低23个百分点);采用m=10的混合选择器可将差距缩小至约12个百分点,且成本仅为全重新验证的46%。在跨任务144项事件的均值分析中,混合预言机约束下原子探针与全重新验证的差距在3个百分点内。据我们所知,原子质量探针是首个为组合式机器人策略中技能更新治理准备就绪的、基于原则的实用基元。