We study whether demonstration-curation metrics that detect defective training episodes also improve the downstream behavior-cloning policy that trains on the curated data. On a contact-rich LIBERO pick-and-place benchmark with a controlled structural defect (early gripper release during the carry phase), we find that the two quantities are sharply decoupled. The metric with the highest defect-detection AUROC (0.804) produces the worst curated policy (13.3% task success), while a metric with a substantially lower AUROC (0.638) produces a policy that nearly matches the oracle trained on ground-truth clean data (90.0% vs. 93.3%). We further show that five of the seven metrics we evaluate exploit episode length as a trivial proxy for the defect label, a confound that inflates reported AUROCs to near-perfect values and disappears once episode length is controlled. Across all conditions, the contaminated baseline succeeds on only 3.3% of rollouts, and the two best curation methods close this to within 3 percentage points of the 93.3% oracle ceiling. Our results argue that curation methods should be evaluated by the policy they produce, not the defects they flag, and that any curation benchmark must control for episode length before reporting detection accuracy. We release the testbed, all metric implementations, and the evaluation pipeline.
翻译:我们研究检测训练片段缺陷的演示文稿筛选指标是否也能改善基于筛选数据训练的下游行为克隆策略。在具有受控结构缺陷(搬运阶段过早释放夹爪)的接触密集型LIBERO拾放基准测试中,我们发现这两个指标之间存在明显脱耦。具有最高缺陷检测AUROC(0.804)的指标生成了最差的筛选策略(任务成功率13.3%),而AUROC较低(0.638)的指标产生的策略几乎与基于真实干净数据训练的最优策略相匹配(90.0%对比93.3%)。我们进一步表明,所评估的七项指标中有五项将片段长度作为缺陷标签的简单代理变量加以利用,这一混杂因素将报告的AUROC值抬高至接近完美的水平,且在控制片段长度后该效应消失。在所有条件下,受污染基线策略仅能在3.3%的试验中成功,而两种最优筛选方法将其提升至与93.3%最优上限相差3个百分点以内。我们的研究结果表明,筛选方法应依据其产生的策略进行评估,而非其标记的缺陷;任何筛选基准测试在报告检测准确性之前必须控制片段长度。我们公开测试平台、所有指标实现及评估流程。