Manipulation demonstrations have temporal phase structure, and a natural hypothesis is that demonstration-curation metrics should be applied within phases rather than globally. The idea is to segment each trajectory into phases, score each phase with the metric that is locally most informative, and then aggregate. This follows directly from prior work showing that a single global metric can be the best detector of a defect and yet the worst curator of the resulting policy. We test the per-phase hypothesis on three contact-rich LIBERO pick-and-place tasks with a controlled early-release structural defect, comparing phase-gated curation against the same metrics applied uniformly and against a strong single global metric. Across all three tasks and five random seeds per condition, phase-gated curation is never the best curation strategy, and it is the worst of the three on two of the three tasks (Task 1: 86.0 vs. 92.0 for global; Task 3: 22.7 vs. 48.0 for uniform). We trace the failure to a concrete mechanism. When the defect signal is concentrated in a single phase, rank-aggregating across phases dilutes that signal with uninformative scores from defect-free phases, selecting a worse demonstration subset than simply applying the defect-informative metric everywhere. We further show that the per-phase metric selection does not transfer across tasks, since no phase shares a winning metric between any two tasks, so the selection cannot be reused and must be re-derived per task from a noisy sweep. These results bound a plausible and previously untested method, and they argue that practitioners should prefer identifying a single defect-informative metric over decomposing curation by phase. We release the full pipeline, all metric implementations, and per-seed results.
翻译:操作演示具有时间阶段结构,自然的假设是演示筛选指标应在各阶段内而非全局应用。其思路是将每条轨迹分割为阶段,用局部信息最丰富的指标对每个阶段打分,然后进行聚合。这直接源于先前研究:单个全局指标可能成为缺陷的最佳检测器,却会导致产生最差的策略筛选结果。我们通过带有受控早期释放结构缺陷的三个高接触性LIBERO拾取-放置任务测试这一逐阶段假设,将阶段门控筛选与统一应用相同指标及单一强全局指标进行对比。在全部三个任务及每个条件五个随机种子的实验中,阶段门控筛选从未成为最优策略,且在三个任务中的两个任务上表现最差(任务1:86.0对比全局92.0;任务3:22.7对比统一48.0)。我们将失败归因于具体机制:当缺陷信号集中于单一阶段时,跨阶段排序聚合会因无缺陷阶段提供的无效信息稀释该信号,导致选取的演示子集劣于在全部阶段应用缺陷信息指标。我们进一步证明,逐阶段指标选择无法跨任务迁移——任何两个任务之间不存在共享最优指标的阶段,因此选择结果无法复用,必须通过噪声扫描为每个任务重新推导。这些结果界定了这种看似合理但未经测试的方法的适用范围,表明实践者应优先识别单一缺陷信息指标而非按阶段分解筛选。我们公开完整流程、所有指标实现及每个种子的结果。