Offline evaluation of agentic systems often collapses trajectories to terminal success, discarding information about partial progress and inducing widespread ties, creating substantial statistical inefficiency by reducing effective sample size and weakening the ability to distinguish systems. We propose preference-based trajectory evaluation, which compares trajectories directly through temporal preferences over progress and time-to-return profiles. We find that, across diverse agentic and interactive benchmarks, standard success-based metrics produce tied comparisons on roughly 75% of instances, whereas trajectory-aware preferences reduce ties to roughly 35%, improving discriminative power, ranking stability, and data efficiency. Our results suggest that benchmark saturation, often attributed to poor data collection or problem difficulty, may also be explained by the choice of evaluation measure.
翻译:智能体系统的离线评估通常将轨迹简化为最终成功状态,丢弃了部分进展信息并导致大量平局,这通过减少有效样本量和削弱系统区分能力造成显著的统计效率低下。我们提出基于偏好的轨迹评估方法,该方法通过比较任务进展与返回时间曲线上的时间偏好直接对比轨迹。研究发现,在多种智能体与交互式基准测试中,基于成功率的传统指标约有75%的实例产生平局比较,而轨迹感知偏好将平局率降至约35%,提升了区分能力、排名稳定性与数据效率。实验结果表明,通常归因于数据收集不足或问题难度的基准测试饱和现象,也可能由评估指标的选择所致。